Skip to main content

Counting Tokens in 7 Data Formats: Pretty JSON Costs 3x CSV

A developer tests 20-row product data across seven formats using the GPT-4o tokenizer, revealing pretty JSON consumes triple the tokens of CSV.

AI-written
Inewgen
03 Oct 2026Source: Dev.to2 min read (0 views)
Share
Counting Tokens in 7 Data Formats: Pretty JSON Costs 3x CSV

Stock photo for illustration only, not from the actual event

Font size
  • A developer tested 20 rows of product data across seven distinct data formats.
  • Pretty-printed JSON consumes approximately three times the tokens of CSV format.
  • YAML features fewer characters than minified JSON yet uses more tokens due to line-by-line key repetition.
  • Dropping unused fields, null values, and extra decimals reduces LLM input costs effectively.

Many developers paste data directly into LLM prompts using JSON formats like JSON.stringify(data, null, 2), rarely checking the underlying token costs. Software developer Jaehyun Cho decided to investigate by taking a single data table, formatting it in seven different ways, and counting the resulting tokens to see the difference.

chromebook notebook computer office desk workspace no logo

Stock photo for illustration only, not from the actual event

The test utilized a product table containing 20 rows and 5 fields including id, name, price, in_stock, and category. A single CSV row looks like 1001, Wireless Mouse, 9.99, false, electronics. Token counts were measured using o200k_base, the tokenizer powering GPT-4o and subsequent OpenAI models, revealing that pretty-printed JSON requires about 3x the tokens of CSV.

3xPretty JSON tokens vs CSV

YAML proved to be an anomaly in the comparison, containing fewer characters than minified JSON while consuming more tokens. This occurs because every field in YAML occupies its own line and retains its own key. Since coding agents frequently transmit large volumes of code, optimization strategies become crucial for scalable applications.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

Understanding tokenization mechanics helps engineers and AI developers optimize system performance and API expenditure. Tokens represent sub-word units rather than direct character counts, meaning compact data serialization directly impacts operational overhead at production scale.

Additional savings can be achieved by stripping away unused fields, null values, and extra decimal places. Attaching a 20-row table to every request—scaling to 1,000 requests daily and 30,000 monthly at $2 per million input tokens—turns minor per-request overhead into a tangible business expense that scales with large tables, RAG results, and API responses.

To assist developers, a free browser-based token counter utilizing the o200k tokenizer has been developed, allowing users to paste data in two formats and compare token counts and model costs locally without uploading sensitive data.

Source: Dev.to

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article