YUAN Benchmarks NVIDIA Cosmos 3 and Qwen 3.5 on Jetson T5000: Advancing Real-Time Edge Reasoning AI
As AI moves from cloud data centers to the physical world, next-generation Edge Video AI requires an escalating demand for multimodal reasoning, low latency, and high-concurrency processing capabilities. To break through the limitations of traditional object detection, YUAN has conducted in-depth performance benchmarks for three next-generation reasoning vision language models (VLM) — NVIDIA Cosmos3-Nano, Cosmos3-Edge, and Qwen3.5-9B — on the NVIDIA Jetson T5000 platform. NVIDIA Cosmos 3 is a world foundation model that can be used as a perception-to-action model or a reasoning vision language model, which we evaluated.
Technical Highlights: The Core Showdown Between Accuracy and Throughput
Taken together, Cosmos3-Nano and Qwen3.5-9B tell a clear throughput story: at 49 concurrent streams, Cosmos3-Nano delivers more than 2x the throughput of Qwen3.5-9B for only about 1% lower accuracy (96.31% vs. 97.44%) — a strong tradeoff for high-density, multi-camera edge deployments.
The benchmark results indicate that both models exhibit exceptional event reasoning capabilities at the edge, each possessing unique architectural advantages:
Qwen3.5-9B | The Ultimate Balance in Semantic Understanding:
Achieved an accuracy of 97.44% and an F1 Score of 91.15. Under the same workload of 49 concurrent streams, it delivered a throughput of 215.37 total tokens/s. It provides an excellent overall balance between precision and recall, making it the premier choice for applications requiring highly reliable semantic understanding and deep contextual reasoning.
Cosmos3-Nano | A Throughput Monster Engineered for Multi-Camera Video Streams:
Achieved an accuracy of 96.31% and an F1 Score of 88.14. Under an extreme workload of 49 concurrent streams, it sustained an impressive throughput of 525.01 total tokens/s, delivering more than twice the throughput performance of Qwen3.5-9B. (Both figures above use vLLM with NVFP4 quantization — YUAN's production-optimized configuration.)
Extending the Family: Cosmos3-Edge on an Equal Footing
To complete the picture, YUAN also benchmarked Cosmos3-Edge — the lightest member of the Cosmos3 family, purpose-built for the most constrained edge devices. Because Cosmos3-Edge's architecture doesn't yet support NVFP4 quantization (via llm-compressor) or vLLM/vLLM-Omni, YUAN benchmarked all three models together in full FP32 precision on the Hugging Face Transformers runtime, giving a clean, apples-to-apples baseline across the whole family, independent of quantization or serving-framework optimizations.

On this equal footing, F1-score — the fairer cross-model metric, since raw accuracy alone can be skewed by class imbalance — scales cleanly with model size: Qwen3.5-9B leads at 90.32, followed by Cosmos3-Nano at 88.32, and Cosmos3-Edge at 87.40. Cosmos3-Edge trades a modest amount of reasoning accuracy for a dramatically smaller footprint.

At single-stream (concurrency = 1) throughput, the pattern flips, as expected for an unquantized, unbatched baseline: Cosmos3-Edge is the fastest at 7.37 tokens/s, versus 3.83 tokens/s for Cosmos3-Nanoand 2.95 tokens/s for Qwen3.5-9B — making Cosmos3-Edge the natural choice when every millisecond and milliwatt counts, such as single-camera or power-constrained edge deployments. For high-concurrency, multi-camera production workloads, Cosmos3-Nano's optimized NVFP4 + vLLM configuration above remains YUAN's recommended path.
A New Era: Moving from Video Analytics to Physical AI
Traditional Video AI relies on predefined object detection, which lacks flexibility. By combining YUAN's edge reasoning technology with multimodal Vision-Language Models (VLMs), operators can now query the system directly using natural language:
“Has a traffic accident occurred on the road?” “Is there an anomaly in the monitored area?” By deeply integrating the powerful computing performance of the NVIDIA Jetson T5000 with YUAN's video capture, multi-camera integration, SmartNVR, and SmartVMS technologies, developers can achieve real-time scene understanding and event reasoning locally. This not only safeguards data privacy but also accelerates the transformation of traditional video analytics into Physical AI and Agentic AI applications with autonomous decision-making capabilities.
More Information
https://blogs.nvidia.com/blog/siggraph-news-2026/#cosmos-3