Counting Tokens in 7 Data Formats: Pretty JSON Costs 3x CSV
A developer tests 20-row product data across seven formats using the GPT-4o tokenizer, revealing pretty JSON consumes triple the tokens of CSV.

Stock photo for illustration only, not from the actual event
- A developer tested 20 rows of product data across seven distinct data formats.
- Pretty-printed JSON consumes approximately three times the tokens of CSV format.
- YAML features fewer characters than minified JSON yet uses more tokens due to line-by-line key repetition.
- Dropping unused fields, null values, and extra decimals reduces LLM input costs effectively.
Many developers paste data directly into LLM prompts using JSON formats like JSON.stringify(data, null, 2), rarely checking the underlying token costs. Software developer Jaehyun Cho decided to investigate by taking a single data table, formatting it in seven different ways, and counting the resulting tokens to see the difference.

Stock photo for illustration only, not from the actual event
The test utilized a product table containing 20 rows and 5 fields including id, name, price, in_stock, and category. A single CSV row looks like 1001, Wireless Mouse, 9.99, false, electronics. Token counts were measured using o200k_base, the tokenizer powering GPT-4o and subsequent OpenAI models, revealing that pretty-printed JSON requires about 3x the tokens of CSV.
YAML proved to be an anomaly in the comparison, containing fewer characters than minified JSON while consuming more tokens. This occurs because every field in YAML occupies its own line and retains its own key. Since coding agents frequently transmit large volumes of code, optimization strategies become crucial for scalable applications.
Understanding tokenization mechanics helps engineers and AI developers optimize system performance and API expenditure. Tokens represent sub-word units rather than direct character counts, meaning compact data serialization directly impacts operational overhead at production scale.
Additional savings can be achieved by stripping away unused fields, null values, and extra decimal places. Attaching a 20-row table to every request—scaling to 1,000 requests daily and 30,000 monthly at $2 per million input tokens—turns minor per-request overhead into a tangible business expense that scales with large tables, RAG results, and API responses.
To assist developers, a free browser-based token counter utilizing the o200k tokenizer has been developed, allowing users to paste data in two formats and compare token counts and model costs locally without uploading sensitive data.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment