Skip to main content

Matryoshka: TypeScript MCP Server for Large Files Without LLM Key

Testing Dmitri Sotnikov's Matryoshka MCP tool with War and Peace and a 10k-line log without any LLM API key.

AI-written
Inewgen
07 Oct 2026Source: Dev.to3 min read (0 views)
Share
Matryoshka: TypeScript MCP Server for Large Files Without LLM Key

Stock photo for illustration only, not from the actual event

Font size
  • Matryoshka is an open-source TypeScript project by Dmitri Sotnikov for handling large documents.
  • It operates via an MCP server, allowing agents to use a query language instead of reading entire files.
  • The test evaluates engine performance using War and Peace and a 10,000-line log without an LLM key.
  • Results show query execution times between 1 and 430 ms with significant token savings.

When a coding agent needs data from a file exceeding its context window, it typically slices the file or relies on a RAG index, both of which often lose information at the seams. Matryoshka introduces a third approach to solve this exact bottleneck.

Built in TypeScript by Dmitri Sotnikov, author of the Luminus Clojure framework, the project draws inspiration from the Recursive Language Models paper by Alex L. Zhang, Tim Kraska, and Omar Khattab from MIT CSAIL, serving as an independent implementation of the concept.

terminal command line server code workspace

Stock photo for illustration only, not from the actual event

As of October 4, 2026, the repository recorded 149 stars, 20 forks, and zero open issues under the Apache-2.0 license, with the latest npm release being matryoshka-rlm 0.2.40 on May 17, 2026. The tester put the framework through its paces on real files while operating under the strict constraint of having no LLM API key and no local Ollama.

97%+Token savings claimed by project
149GitHub repository stars
2.5sPackage installation time

Rather than writing JavaScript or Python, the model issues commands in Nucleus, a compact S-expression language such as (grep "ERROR") and (count RESULTS). The Lattice engine parses and executes these commands against the document, storing results in an in-memory SQLite database while returning concise stubs to the client.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

"The README claims this gives '97%+ token savings', and '80%+ token savings compared to reading files directly' for the MCP server."

Dmitri Sotnikov

Tested on a Linux x86_64 machine equipped with 8 vCPUs and 15 GB of RAM, local installation finished in 2.5 seconds, consuming 220 megabytes within node_modules. Crucial operational caveats included avoiding plain npx lattice-mcp commands to prevent downloading unrelated packages, recommending the bundled binary instead.

Enabling an agent to query server-side documents using a specialized query language and returning only lightweight handles or previews fundamentally reduces the overhead of feeding massive text blocks into LLM context windows. This architecture highlights an efficient paradigm for managing large codebases, literary works, or extensive system logs without overwhelming token limits.

Source: Dev.to

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article