• Menu Item
  • STORIES
    • TECH
    • AUTOMOTIVE
    • GUIDES
    • OPINIONS
  • REVIEWS
    • READERS’ CHOICE
    • ALL REVIEWS
    • ━
    • SMARTPHONES
    • CARS
    • HEADPHONES
    • ACCESSORIES
    • LAPTOPS
    • TABLETS
    • WEARABLES
    • SPEAKERS
    • APPS
  • FINAL CUT
    • TV & MOVIES REVIEWS
    • SPOTLIGHT
  • GAMING
    • GAMING NEWS
    • GAMING REVIEWS
  • n
  • m
  • menu
  • SEARCH
  • STORIES
    • TECH
    • AUTOMOTIVE
    • GUIDES
    • OPINIONS
  • REVIEWS
    • READERS’ CHOICE
    • ALL REVIEWS
    • ━
    • SMARTPHONES
    • CARS
    • HEADPHONES
    • ACCESSORIES
    • LAPTOPS
    • TABLETS
    • WEARABLES
    • SPEAKERS
    • APPS
  • FINAL CUT
    • TV & MOVIES REVIEWS
    • SPOTLIGHT
  • GAMING
    • GAMING NEWS
    • GAMING REVIEWS
  • n
  • m
Follow US

NVIDIA targets faster AI agents with Nemotron 3.5 Lightning

ADAM D.
ADAM D.
2 minutes ago

NVIDIA is expanding its ambitions beyond the hardware that powers artificial intelligence, introducing a new open model and routing technology aimed at developers building AI agents that need to operate continuously, efficiently and across multiple specialised tasks.

The centrepiece is Nemotron 3.5 Lightning, a lightweight and customisable open model designed for high-volume agentic AI workloads. Rather than positioning a single large model as the answer to every request, NVIDIA is leaning into an increasingly common approach in AI development: using several specialised models together and selecting the right one for each stage of a task.

NVIDIA says Nemotron 3.5 Lightning can generate output at up to four times the speed of comparable models in its class, resulting in agentic tasks being completed up to 30% faster. Those figures are NVIDIA’s own performance claims, so real-world results will inevitably depend on the workload, hardware and way developers configure their agents.

Speed is only part of the problem NVIDIA is trying to address. Running autonomous or always-on AI agents can quickly become expensive when every step is sent through a large, computationally demanding model. That is where NeMo Switchyard enters the picture.

The open-source routing library is designed to examine individual stages of an agent workflow and direct them toward the model considered most appropriate for the job. In principle, straightforward requests could be handled by smaller and faster models, while more complicated reasoning is passed to more capable systems. The goal is to give developers another way of balancing response quality, latency and inference costs without manually deciding which model handles every operation.

This multi-model approach could become increasingly important as AI agents move beyond relatively simple chatbot interactions. Agents working across coding, cybersecurity, research or enterprise software may need to perform dozens of individual operations to complete a single objective. Using the most powerful available model for every one of those steps can be unnecessarily expensive and slow.

NVIDIA says several companies are already customising Nemotron 3.5 Lightning for their own applications. The list includes CrowdStrike in cybersecurity, Harvey in legal AI, CodeRabbit in software development, Fastino and Lila Sciences, with potential applications spanning areas including finance and healthcare.

The release also fits NVIDIA’s broader effort to establish a larger presence in open AI development. The company has backed initiatives supporting open-weight models and has been involved in the Open Secure AI Alliance, while its Nemotron family provides an alternative to the increasingly crowded selection of open and openly available models from competing AI developers.

For NVIDIA, there is an obvious strategic benefit. More capable open models and agent frameworks can encourage more AI workloads, and those workloads ultimately need computing infrastructure. But the more interesting development for AI builders may be the shift away from assuming one model should do everything. If agentic AI is going to operate continuously at scale, choosing when not to use the biggest model could become just as important as choosing which model to use.

Share
What do you think?
Happy0
Sad0
Love0
Surprise0
Cry0
Angry0
Dead0

WHAT'S HOT ❰

Samsung Galaxy Buds receive FDA clearance for hearing aid feature
UAE becomes third OpenAI region to support inference residency
UAE viewers can watch tonight’s total solar eclipse live on Twitch
Grok Bot brings persistent AI agents to macOS and iOS
ChatGPT finally lands on Linux with Codex and Work support
Follow US
Caffeine. Chaos. Content.
© Absolute Geeks Media FZE LLC 2014–2026.
Proudly made in Dubai, UAE ❤️
Contact · About · Editorial Policy · Privacy Policy
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?