Skip to main content

Prime Intellect Launches Prime Inference for Frontier Models

Prime Intellect unveils Prime Inference, a serverless and reserved serving layer for frontier open models optimized for agentic workloads.

AI-written
Inewgen
03 Oct 2026Source: MarkTechPost3 min read (0 views)
Share
Prime Intellect Launches Prime Inference for Frontier Models

Stock photo for illustration only, not from the actual event

Font size
  • Prime Intellect launches Prime Inference serving layer for open artificial intelligence models.
  • Combines NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer technologies.
  • Engineered specifically to handle complex and heavy agentic workloads.
  • Achieves nearly 40 percent lower p90 inter-token latency in rigorous testing.

Prime Intellect has officially introduced Prime Inference, a brand-new serving layer designed to complete the company's open training infrastructure stack. Previously, the firm delivered various post-training tools including prime-rl, verifiers, and sandboxes. Introducing this serving capability closes the loop, allowing deployed models to generate production traces that feed directly back into ongoing training cycles.

Regarding operational performance, the company highlights that its GLM-5.3 endpoint ranks among the fastest available on OpenRouter. Furthermore, it boasts a near-zero tool-call error rate alongside an unbroken 100 percent uptime record since its initial launch.

40%Lower p90 latency
100Tokens per second per GPU

The underlying infrastructure stack integrates several advanced technologies, namely NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer. Developed alongside Inferact and NVIDIA, the team also contributes fixes back upstream. The primary workload target for this infrastructure is agentic applications, where a typical agent turn appends approximately 6,000 tokens onto a 140,000-token prompt. Benchmarks were conducted using SemiAnalysis AgentX combined with injected cold arrivals.

artificial intelligence server hardware rack

Stock photo for illustration only, not from the actual event

A core architectural feature is prefill and decode disaggregation, where prefill and decode operations execute across entirely separate GPU groups. Dynamo manages request routing while vLLM runs the model on each respective group. Decoders pull computed key-value pairs through NIXL, which allowed Prime to report a nearly 40 percent reduction in p90 inter-token latency during testing.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

"Prefill and decode run on separate GPU groups. Dynamo handles routing, and vLLM runs the model on each group."

Prime Intellect Engineering Team

Additionally, the architecture utilizes cache-aware routing. Dynamo's KV-aware router weighs cached prefix overlap against queued work to ensure sessions remain bound to the same decoder between turns. Meanwhile, Mooncake introduces a secondary KV caching tier hosted within host DRAM to further optimize throughput.

Disaggregating prefill and decode phases represents a vital architectural strategy for modern large language models, as both phases impose distinct computational and memory bandwidth demands. Intelligent request routing and advanced KV cache management effectively eliminate bottlenecks, enabling high concurrency and sustained responsiveness for interactive AI workloads.

The targeted interactivity benchmark was set at 100 end-to-end tokens per second per user. At this operational threshold, a 1:4 prefill-to-decode ratio served the highest volume of users, reaching 66 sessions per prefill group at 101 tokens per second per user and 100 output tokens per second per GPU.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article