Paste a fairly ordinary JSON export into a prompt, a few thousand support tickets, say, or a week of user activity logs, and watch what happens to the token count. A file that's maybe 400KB on disk turns into a number that looks like a typo. Somewhere in the back of your mind, a small alarm goes off: that can't be right.
It's right; nothing about your data changed. What changed is that a language model doesn't read JSON the way your code editor does, and the format has been quietly working against you on every single request you've ever sent.
Here's the important part: the usual fix getting a model with a bigger context window barely helps. A million-token window will still burn through your budget re-reading the word "email" ten thousand times, and it will still lose track of the one record buried in the middle of all that repetition. The real issue was never size. It's shape.
The good news is that none of this is unsolvable because developers have been quietly working around JSON's inefficiencies for a while now, and the fixes range from a five-minute tweak (strip the whitespace) to more deliberate architectural choices (rethink how you chunk and retrieve data in the first place). This guide walks through the whole toolkit, in the order you'd actually reach for it.
Why JSON and LLMs Don't Get Along by Default
Before fixing anything, let’s try to understand why the problem exists and how we solve this in JSON.
An LLM doesn't see your JSON the way you do, like you see the structure of a user's array, a role field, a nested address object. However, the model sees tokens: fragments of text that get counted, priced, and packed into a fixed-size context window. And JSON, by design, is extremely generous with characters that carry zero meaning for the model but still cost tokens.
Take a simple array of user records and note every single object repeats the same key names like id, name, email, role; over and over, once per record. On a 10-record array, that's mildly annoying because of repetitive points. On a 10,000-record array, you're paying to re-transmit the word "email" ten thousand times. Add in braces, colons, commas, and quotation marks around every key and string value, and a huge share of your token budget goes to punctuation the model has already seen a thousand times before, not to actual information.
A deeply nested JSON tree can end up spending more tokens on scaffolding than on data because every level of depth adds another layer of brackets and, often, another repeated key structure.
The two cost analyses in JSON are financial and reasoning cost :
- A financial cost: Most LLM APIs charge per token, in and out. Bloated JSON means you're paying for punctuation.
- A reasoning cost: This one surprises people; models don't just get slower with more tokens; they get less accurate because relevant details buried in the middle get less attention than details near the start or end. Feed a model 50,000 tokens of repetitive, noisy JSON, and you're not just spending more money; you're making it statistically more likely the model glosses over the one record that actually mattered.
Once you see it this way, "optimizing JSON" stops being a nice-to-have and starts looking like a basic requirement for anything built on top of an LLM at scale.
What "Large" Actually Means to a Model
One thing worth clearing up early: file size and token count aren't the same, and confusing them leads people to either panic unnecessarily or underestimate a real problem.
A 2MB JSON file with lots of numeric data might tokenize into far fewer tokens than a 500KB file full of long descriptive text, because tokenizers work on the density of distinct sub-word chunks, not raw bytes. If you want to know whether your dataset is actually "large" in the way that matters, run it through a tokenizer, not a file size checker.
Context windows across major providers have grown substantially over the past couple of years, with many production models now supporting anywhere from roughly 128,000 tokens up into the millions. That should make this whole conversation obsolete. It doesn't, for three reasons:
- Cost still scales linearly (or worse) with tokens. A bigger window doesn't mean a cheaper one.
- Attention doesn't scale as cleanly as the window does. Independent long-context benchmarks routinely show accuracy dropping well before a model hits its advertised limit, especially on tasks that require pulling a specific fact out of a large, repetitive structure.
- Latency grows with input size. If you're running this inside an agent loop or a production API, every extra thousand tokens is extra time your user is waiting.
So even with a generous context window, the goal isn't "does it fit?", it's "does the model actually use it well, at a cost I'm comfortable with?"
Diagnose Before You Optimize
It's tempting to jump straight to compression tricks, but a quick audit saves you from optimizing the wrong thing. Before touching your dataset, ask three questions:
- How many tokens per record am I actually sending? Run a sample through a tokenizer and calculate a rough tokens-per-record figure. This becomes your baseline.
- What's the actual task? A one-off analysis of a single file is a very different problem from a production agent that fetches JSON from a tool call on every user turn, or a RAG pipeline indexing a hundred thousand records.
- Does the model need the whole dataset, or just an answer derived from it?
With that baseline in hand, the fixes below stop being guesswork and start being measurable improvements.
The Core Optimization Playbook
1. Strip What the Model Never Needed
Before any structural changes, look at what's actually in the payload. Production JSON is often full of fields that exist for internal tooling, not the task at hand: debug flags, internal IDs nobody asked about, timestamps down to the millisecond, null fields, empty arrays, and verbose metadata wrappers. None of it helps the model, and all of it costs tokens.
A quick pruning pass, even a manual one, routinely removes 10 to 30 percent of a payload before you've touched formatting at all. If you're building a pipeline, this is a great first automated step: define a schema of "fields the model is allowed to see" and filter everything else out before the data ever reaches your prompt.
2. Minify, Without Hesitation
Pretty-printed JSON, with its indentation and line breaks, exists for human readers. Models don't care about it, and every space and newline character is still a token you're paying for. Minifying stripping whitespace down to the bare structural minimum is close to a free win: no information is lost, and the token count drops immediately.
The only caveat: keep a pretty-printed version around for your own debugging. Minified JSON is unreadable at 2 am when something breaks.
3. Flatten Nested Structures Where You Can
Deep nesting means deep repetition of brackets and, frequently, of key paths. If your data has three or four levels of nesting purely for organizational reasons, not because the relationships are genuinely hierarchical, flattening it into a shallower structure (using dot-notation keys like address. city instead of a nested address object) can meaningfully cut overhead while keeping every value intact.
That's not always the right call, though. If the nesting reflects a real one-to-many relationship an order with multiple line items, a user with multiple sessions forcing it flat can actually confuse the model about which values belong together. Flatten what's structurally shallow. Leave genuinely hierarchical relationships nested, but make sure they're nested efficiently (see the next point).
4. Switch Formats: JSON Isn't the Only Option
Here's the part most people miss entirely: the model doesn't care what format your data arrives in; it only cares about extracting meaning from tokens. That opens the door to formats specifically designed to represent structured data with far less overhead.
For arrays of uniform objects, the single most common pattern in LLM prompts is a format called TOON (Token-Oriented Object Notation), which tends to be the biggest single win available. TOON combines the readability of YAML's indentation with the compactness of CSV-style tabular rows. Instead of repeating every key for every object in an array, it declares the field names once and lists the values in a simple, comma-separated row per record.
That's the practical rule of thumb for when to reach for a format like TOON:
- Use it for: arrays of objects, catalog data, user lists, log entries, tabular exports, anything with repeated keys across many records.
- Skip it for: small payloads (a handful of fields, not worth the conversion step), deeply irregular or inconsistently shaped JSON, or situations where you're calling an API that strictly requires standard JSON on the way out.
And because the conversion is fully reversible, you're not locking yourself into a new format; you can convert JSON to TOON for the prompt, get a TOON-shaped response back, and convert it straight back to JSON for your application code. If you want to see this on your own data before committing to it, the JSON to TOON converter runs entirely in your browser, shows you the token count on both sides in real time, and never sends your data anywhere. It's worth pasting in a real sample from your project to see the actual number, rather than trusting a general estimate.
For data that's genuinely tabular with no nesting at all, plain CSV is worth considering too. For configuration-style data that a human will also be reading, YAML strikes a nice middle ground. But for the specific case that trips up most LLM pipelines, big arrays of structured records, TOON is purpose-built for exactly that problem, which is a large part of why it's gained traction quickly among people building RAG systems and agent pipelines.
5. Shorten Keys, Carefully
Long, descriptive key names like transaction_timestamp_utc are considerate to human developers and expensive for token budgets. One option is to define short aliases, ts instead of timestamp, qty instead of quantity, and declare the mapping once, either in your system prompt or as a small header at the top of the data itself. Current-generation models are quite good at holding onto a defined mapping and applying it consistently.
Use this one with some restraint. Over-abbreviate and you risk ambiguity (does st mean state or status?), which can cost you more in incorrect outputs than it saved in tokens. Keep abbreviations short but unambiguous, and always define them explicitly rather than assuming the model will guess correctly.
6. Chunk by Structure, Not by Character Count
If your dataset genuinely doesn't fit in one request, you'll need to split it. This is where a lot of pipelines quietly break, because splitting a large JSON blob by raw character count is convenient to code and terrible in practice; it has no respect for object boundaries. Cut a record in half, and you've handed the model a malformed fragment it has to either discard or hallucinate around.
Chunk by structure instead. If you're working with an array, split along record boundaries; every chunk should contain complete, valid objects, never a partial one. Tools like jq or a JSONPath library make this straightforward to script.
The trickier issue shows up with nested data: split a large nested tree into pieces, and each fragment loses the parent context that gave it meaning. A department object without the company it belongs to is just floating information. The fix is to carry a lightweight breadcrumb with every chunk, a short header noting which parent path this fragment came from, so the model always knows where a piece fits in the bigger structure, even though it's only seeing a slice of it. It costs a handful of extra tokens per chunk and saves you from a model that starts inventing relationships to fill the gap.
7. Filter and Query Before You Send Anything
This is the highest-leverage step on this entire list, and it's the one most people reach for last instead of first: don't send the whole dataset if the task only needs a slice of it.
If you're analyzing a week of metrics but the question is about Tuesday, filter to Tuesday before the data touches a prompt. If you're answering a question about one customer, query for that customer's record rather than shipping your entire customer table and asking the model to find it. A simple jq filter, a SQL query, or a JSONPath expression run before the prompt is built will usually save more tokens than every formatting trick combined, because it addresses the actual redundancy: sending information the task doesn't need at all.
This also tends to be the fix for time-series or multi-day data specifically. Rather than dumping seven days of raw, granular records into a prompt and hoping the model spots the trend, pre-aggregate what can be aggregated: daily totals, averages, notable outliers, and only include the raw granular rows for the specific window the question actually concerns.
8. Stream and Paginate for Very Large Files
When you're working with files too large to load into memory comfortably, let alone into a prompt, loading the whole thing to trim it down defeats the purpose. Streaming parsers process a file incrementally, one record at a time, without holding the entire structure in RAM, useful both for your own preprocessing scripts and for keeping memory usage sane in production.
If you're building an agent that calls tools through MCP or a similar protocol, and one of those tools can return a large JSON payload, don't let the raw response flow straight into the model's context. Have the tool paginate, summarize, or filter its own output before returning it, and give the agent a way to ask for the next page or a more specific slice if it needs one. A tool that dumps ten thousand rows into context because someone called it without a filter is one of the fastest ways to blow through a token budget in an automated pipeline, and it's entirely preventable at the source.
9. Reach for RAG When You Have a Knowledge Base, Not Just a Payload
There's a meaningful difference between "I need to send this JSON to the model" and "I have a large collection of JSON records and need the model to answer questions about them occasionally."
If you're sitting on thousands of JSON records, a product catalog, a support ticket history, a knowledge base, the right move is usually to embed each record (or a logically coherent group of records) into a vector store and retrieve only the handful that are relevant to a given query. The model never sees the full dataset; it sees a small, relevant slice pulled in on demand. This is the standard retrieval-augmented generation pattern, and it scales to datasets that would never fit in a context window no matter how aggressively you compressed them.
10. Trim Numeric and Repetitive Noise
Two smaller techniques round out the list, and they matter more than their size suggests once you're operating at volume:
- Decimal precision. A price field storing 299.999999999 instead of 300 isn't adding information for most use cases; it's adding tokens. Round to the precision your task actually requires.
- Delta encoding for logs and time series. If you're sending sequential snapshots of the same object, sending the full state every single time is redundant. Sending only what changed between snapshots keeps every update compact without losing any information across the sequence.
Neither of these will transform a bloated dataset on their own, but stacked with the earlier steps, they close the gap between "meaningfully smaller" and "as lean as it's going to get."
Bringing It Together: A Practical Workflow
Read individually, these techniques can feel like a long checklist. In practice, they layer cleanly into a short sequence you can apply to almost any pipeline:
- Audit your baseline token count on a representative sample.
- Strip fields the model doesn't need.
- Filter and pre-aggregate down to what the task actually requires.
- Flatten shallow nesting; keep genuine hierarchies intact.
- Convert repetitive, tabular structures to a token-efficient format like TOON.
- Chunk by structure if it still doesn't fit, with parent-context breadcrumbs attached.
- Re-measure your token count and, just as importantly, test output quality against your original baseline.
That last step gets skipped constantly, and it's the one that matters most. A technique that saves 40 percent of your tokens but quietly degrades accuracy isn't an optimization; it's a trade you haven't measured yet. Always compare outputs before and after, not just token counts.
Mistakes Worth Avoiding
A few patterns show up again and again in pipelines that "optimized" their way into new problems:
- Chunking without preserving relationships. A record split from its parent context is a record the model has to guess about.
- Over-abbreviating keys until they're ambiguous. A few saved tokens aren't worth an incorrect answer.
- Optimizing a format that needs to stay standard JSON on the way out. If you're calling a strict API contract, minify and filter, but don't reformat what needs to stay JSON-shaped.
- Forgetting the round trip. If your workflow needs data back in its original shape, make sure whatever transformation you apply going in is reversible coming out.
Most teams get the bulk of the benefit from three moves: filtering out what the task doesn't need, minifying what's left, and converting genuinely repetitive structures, arrays of similar objects, the single most common shape in real-world JSON, into a format built for token efficiency instead of general-purpose parsing.
If you want to see how much of a difference that last step makes on your own data, run a real sample through the free JSON to TOON converter and compare the token counts side by side. For a lot of teams, it's the single change that turns an expensive, sluggish JSON-heavy prompt into something fast, cheap, and considerably easier for the model to actually reason about.
Frequently Asked Questions
I need to feed a genuinely large dataset into an LLM, what's actually the right first move?
Start by asking whether the model needs the whole dataset at all, or just an answer derived from it. Most of the time it's the latter. Filter or pre-aggregate down to what the current task requires, prune unnecessary fields, and only then worry about format and chunking. Trying to compress an entire dataset before asking whether all of it is needed is the most common wasted effort in this space.
I have several days' worth of JSON data, logs, metrics, whatever, and the model seems to skip over or truncate parts of it when I paste it in. What's going wrong?
This is usually a volume problem hiding as a comprehension problem. Rather than pasting raw, granular records for every day, pre-aggregate what can be summarized (daily totals, trends, anomalies) and only include full granular detail for the specific window your question actually concerns. A model reasoning over a tight, relevant dataset consistently outperforms one wading through a week of raw noise looking for the same answer.
How do I send deeply nested JSON to a model without losing the parent-child relationships once I have to chunk it?
Chunk along natural structural boundaries, never mid-object, and attach a short breadcrumb to each chunk noting where it sits in the original hierarchy. Losing hierarchy during chunking is almost always a byproduct of splitting by character count instead of by structure; fix the splitting method and the relationship problem usually goes away with it.
Will converting JSON to TOON or a similar format lose any information?
No, a properly implemented conversion is fully reversible. The transformation changes how the data is represented (removing repeated keys and unnecessary punctuation), not what the data contains. You can convert back to standard JSON at any point and get the original structure and values intact, which is why it's safe to use TOON purely as a prompt-time optimization while keeping standard JSON everywhere else in your application.
