Pokee AI Releases Pokee-Isaac 28B with 10M-Token Context Window
Pokee AI launches the Pokee-Isaac 28B agentic model featuring a 10-million token context window, designed for secure deployment within customer boundaries and on-device hardware.

Stock photo for illustration only, not from the actual event
- Pokee AI introduces the Pokee-Isaac model with 28 billion parameters
- Supports a massive 10-million token context length accurately
- Available via API and licensed for VPC and on-premises deployment
- Outperforms competing models across major benchmark evaluations
The artificial intelligence landscape welcomes Pokee-Isaac 28B, an advanced agentic model developed by Pokee AI designed to handle an enormous context window of 10 million tokens. The model is offered as a licensed solution through an OpenAI-compatible developer API, with deployment options tailored for virtual private clouds (VPCs), on-premises environments, or local hardware devices.
Regarding infrastructure support, the launch announcement details Day-0 compatibility with vLLM and SGLang, alongside single-GPU serving requirements beginning with an RTX 4090 or equivalent hardware. However, the research team published measurements derived exclusively from a B200-class GPU, meaning consumer-GPU deployment claims should be viewed as vendor guidance rather than benchmarked outcomes.

Stock photo for illustration only, not from the actual event
When evaluated across rigorous benchmark suites, the Pokee-Isaac model delivers exceptional performance metrics. According to the RULER evaluation framework, the model maintains an accuracy score remaining above 93.3% across every tested length, concluding at 93.3% at the 10-million token mark. By comparison, baseline models such as GPT-5.6 Luna and Gemini 3.5 Flash Lite track closely up to 512K before encountering context-overflow errors at 1M.
Furthermore, in the MRCR v2 evaluation utilizing 8 needles, Isaac achieves scores of 0.607, 0.743, and 0.500 at context lengths of 256K, 512K, and 1M respectively, expanding its performance margin over Gemini from 0.133 to 0.295 throughout the evaluation sweep. In DTAP red-teaming assessments, Isaac records the lowest direct attack success rate at 36.0, indirect at 35.2, and combined at 35.6, accompanied by an 82.5% benign success rate, noting that baseline models executed under the stock runner while Isaac utilized the Pokee harness.
The capability of modern artificial intelligence models to process a 10-million token context window entirely within customer boundaries represents a major leap forward for data security and privacy. Enabling local execution within private enterprise infrastructure or directly on-device significantly mitigates the compliance risks associated with transmitting sensitive proprietary data to external cloud servers, catering directly to heavily regulated sectors like finance and healthcare.
Analyzing performance characteristics on a single B200-class GPU under the RULER workload, time to first token (TTFT) is measured at 23.6 seconds at 1M context length and 72.9 seconds at 10M. Meanwhile, prefill throughput scales upward alongside context expansion, shifting from 42,400 to 137,200 tokens per second, indicating that a tenfold increase in prompt length incurs roughly three times the TTFT overhead.
Provisional list pricing is established at $0.15 per million input tokens and $1.00 per million output tokens. Additionally, Isaac operates fully on-device across hardware architectures including Intel Arc Pro B70, Core Ultra Series 3 (Panther Lake), and Qualcomm Snapdragon X2 Elite processors.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment