Matryoshka: TypeScript MCP Server for Large Files Without LLM Key
Testing Dmitri Sotnikov's Matryoshka MCP tool with War and Peace and a 10k-line log without any LLM API key.

Stock photo for illustration only, not from the actual event
- Matryoshka is an open-source TypeScript project by Dmitri Sotnikov for handling large documents.
- It operates via an MCP server, allowing agents to use a query language instead of reading entire files.
- The test evaluates engine performance using War and Peace and a 10,000-line log without an LLM key.
- Results show query execution times between 1 and 430 ms with significant token savings.
When a coding agent needs data from a file exceeding its context window, it typically slices the file or relies on a RAG index, both of which often lose information at the seams. Matryoshka introduces a third approach to solve this exact bottleneck.
Built in TypeScript by Dmitri Sotnikov, author of the Luminus Clojure framework, the project draws inspiration from the Recursive Language Models paper by Alex L. Zhang, Tim Kraska, and Omar Khattab from MIT CSAIL, serving as an independent implementation of the concept.

Stock photo for illustration only, not from the actual event
As of October 4, 2026, the repository recorded 149 stars, 20 forks, and zero open issues under the Apache-2.0 license, with the latest npm release being matryoshka-rlm 0.2.40 on May 17, 2026. The tester put the framework through its paces on real files while operating under the strict constraint of having no LLM API key and no local Ollama.
Rather than writing JavaScript or Python, the model issues commands in Nucleus, a compact S-expression language such as (grep "ERROR") and (count RESULTS). The Lattice engine parses and executes these commands against the document, storing results in an in-memory SQLite database while returning concise stubs to the client.
"The README claims this gives '97%+ token savings', and '80%+ token savings compared to reading files directly' for the MCP server."
Dmitri Sotnikov
Tested on a Linux x86_64 machine equipped with 8 vCPUs and 15 GB of RAM, local installation finished in 2.5 seconds, consuming 220 megabytes within node_modules. Crucial operational caveats included avoiding plain npx lattice-mcp commands to prevent downloading unrelated packages, recommending the bundled binary instead.
Enabling an agent to query server-side documents using a specialized query language and returning only lightweight handles or previews fundamentally reduces the overhead of feeding massive text blocks into LLM context windows. This architecture highlights an efficient paradigm for managing large codebases, literary works, or extensive system logs without overwhelming token limits.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment