Cost saved
53.7% lower$2.47$4.59 → $2.12 · control → GPTZIPCompress the tokens
The only
lossless token compression.
It's also the fastest.
Before GPTZIP
182
tokens
After
129
tokens
−29.1%
input reduction
27 ms
Compression Time
AI savings calculator
What could fewer tokens save your business?
Pick the work your team does and enter your monthly AI API spend. We'll estimate the impact of sending smaller inputs.
Which AI workflow are you paying for?
Enter the input-token portion of your usage-based API spend. Do not include seats or subscriptions.
Your potential savings
A smaller input bill, every month.
Enter your monthly spend to see your estimate.
Your result updates instantly.
Projection uses the 19.9% median token reduction found in the closest landing-page benchmark workloads: CoQA · SQuAD v2 · GPQA Diamond · FinanceBench · HumanEval. Output-token charges, subscriptions, seats, and other fees are not reduced.
Benchmark data
LLM accuracy increases after GPTZIP compression.
Accuracy and compression
Methods
Datasets
| Benchmark | Method | Median tokens removed | Accuracy change |
|---|---|---|---|
| FinanceBench | GPTZIP | 19.4% | +6.6pp |
| FinanceBench | Headroom | 16.1% | −12.5pp |
| FinanceBench | Bear (TheTokenCompany) | 1.9% | +4.1pp |
| FinanceBench | LLM-Lingua (MSFT) | 59.2% | −1.6pp |
| SQuAD v2 | GPTZIP | 32.8% | +17.2pp |
| SQuAD v2 | Headroom | 10.8% | +24.7pp |
| SQuAD v2 | Bear (TheTokenCompany) | 9.3% | +34.3pp |
| SQuAD v2 | LLM-Lingua (MSFT) | 61.0% | +2.8pp |
| CoQA | GPTZIP | 42.6% | +12.0pp |
| CoQA | Headroom | 51.8% | −9.7pp |
| CoQA | Bear (TheTokenCompany) | 13.1% | +16.4pp |
| CoQA | LLM-Lingua (MSFT) | 62.8% | −7.6pp |
| GPQA Diamond | GPTZIP | 19.9% | +5.2pp |
| GPQA Diamond | Headroom | 10.6% | +13.8pp |
| GPQA Diamond | Bear (TheTokenCompany) | 4.9% | +37.9pp |
| GPQA Diamond | LLM-Lingua (MSFT) | 48.4% | +4.0pp |
Method comparison
GPTZIP is the only lossless compression across use-cases.
| Benchmark | Uncompressed | GPTZIP | Headroom | Bear (TheTokenCompany) | LLM-Lingua (MSFT) |
|---|---|---|---|---|---|
| FinanceBench | 86.7%Accuracy | +6.6ppLossless 93.3% accuracy18.8% compression rate19.4% median40.1% max | −12.5pp 74.2% accuracy16.1% median83.8% max | +4.1ppLossless 90.8% accuracy1.9% median14.1% max | −1.6pp 85.1% accuracy59.2% median68.3% max |
| LongBench v2 | N/AAccuracy | Not available | Not available | Not available | Not available |
| SQuAD v2 | 56.1%Accuracy | +17.2ppLossless 73.3% accuracy33.5% compression rate32.8% median41.4% max | +24.7ppLossless 80.8% accuracy10.8% median51.1% max | +34.3ppLossless 90.4% accuracy9.3% median23.1% max | +2.8ppLossless 58.9% accuracy61.0% median67.6% max |
| CoQA | 76.8%Accuracy | +12.0ppLossless 88.8% accuracy41.8% compression rate42.6% median55.2% max | −9.7pp 67.1% accuracy51.8% median77.3% max | +16.4ppLossless 93.2% accuracy13.1% median25.9% max | −7.6pp 69.2% accuracy62.8% median66.0% max |
| GPQA Diamond | 29.3%Accuracy | +5.2ppLossless 34.5% accuracy22.5% compression rate19.9% median76.1% max | +13.8ppLossless 43.1% accuracy10.6% median95.4% max | +37.9ppLossless 67.2% accuracy4.9% median19.8% max | +4.0ppLossless 33.3% accuracy48.4% median78.9% max |
Workload fit
Compression depends on the work you send.
Long contexts and short questions do not compress the same way. These rows use the closest measured public benchmark instead of one number for every workload.
| Workload | Public proxy | Median removed | Accuracy change |
|---|---|---|---|
| Financial and regulatory analysis | FinanceBench | 19.4% | +6.6pp |
| Conversational and extractive QA | CoQA · SQuAD v2 | 32.8–42.6% | +12.0pp to +17.2pp |
| Expert reasoning | GPQA Diamond | 19.9% | +5.2pp |
Compression speed
330-token measured prompt
Terminal-Bench 2.1
Lossless task outcomes with measured cost savings.
GPTZIP keeps the verifier outcome on the shown task pairs while using fewer billed tokens and dollars. The view combines task resolution, cost per successful task, turns, and cache tiers so the compression result is auditable.
The tasks displayed were chosen at a random draw, more task results are available upon request.
Tokens saved
67.6% lower6.47M9.57M → 3.10M · control → GPTZIPBilled turns
57.0% lower85149 → 64 · control → GPTZIPQuality held
100%3/3 shown pairs3 of 5 paired tasks shownOutcome economics
Cost per successful task
Successful task pairs
Results by task and programming language
| Task | Programming language | Billed turns | Total tokens | Cost | Cache read |
|---|---|---|---|---|---|
write-compressorCreate a compressed payload accepted by an existing C decompressor within the size limit. | Any (Rust reference) | 39 → 1171.8% lower | 1.81M → 294.7k83.7% lower | $1.27 → $0.3373.8% lower | 1.69M → 257.3k84.8% lower |
bn-fit-modifyRecover a Bayesian-network DAG, intervene on Y, and produce the requested sample files. | R-oriented | 17 → 1041.2% lower | 288.2k → 168.1k41.7% lower | $0.28 → $0.2413.7% lower | 251.6k → 128.5k48.9% lower |
build-cython-extBuild pyknotid Cython extensions and make the package work with the installed NumPy version. | Python/Cython | 93 → 4353.8% lower | 7.47M → 2.64M64.7% lower | $3.05 → $1.5549.0% lower | 7.27M → 2.46M66.1% lower |
Private prompt compression
Compress private prompts safely.
The masking method runs on your browser, computer, or server. GPTZIP gets the masked text—not the original text. Drag the slider to see the live example.
Original input: Customer support case for Aurora Health: Maya Chen reported a duplicate charge on order 483901. The support agent must verify the billing event, preserve the refund policy, and answer whether the customer can keep the replacement device. Customer support case for Aurora Health: Maya Chen reported a duplicate charge on order 483901. The support agent must verify the billing event, preserve the refund policy, and answer whether the customer can keep the replacement device. Masked text: Client help type in Dawn Wellness: Alice David covered some twin care at club 145302. Any help factor need control each request issue, save each payment insurance, but reply house each buyer may hold each switch twist. Client help type in Dawn Wellness: Alice David covered some twin care at club 145302. Any help factor need control each request issue, save each payment insurance, but reply house each buyer may hold each switch twist.
Drag to compare original and masked text.
What leaves your device
The original data does not go to GPTZIP.
The customer-controlled device keeps the original text and the information needed to restore it. Only masked text goes to the GPTZIP service.
Your browser
The original prompt and identifiers stay on your device.
GPTZIP server
Receives only masked text and returns compressed masked output.
Compressed output
A shorter masked request comes back to your device.
Restore in the browser
The browser puts the original terms back into the user-facing text.
88
masked input
68
sent request
20
tokens saved
Customer support case for Aurora Health: Maya Chen reported a duplicate charge on order 483901. The support agent must verify the billing event, preserve the refund policy, and answer whether the customer can keep the replacement device. Customer support case for Aurora Health: Maya Chen reported a duplicate charge on order 483901. The support agent must verify the billing event, preserve the refund policy, and answer whether the customer can keep the replacement device.For developers
Less tokens. More LLM quality.
Add GPTZIP before the LLM call or use GPTZIP Proxy. You do not need to change models or rebuild your harness.
- One API step for apps you own
- OpenAI-compatible proxy path for supported clients
- Before-and-after data you can inspect
For enterprise
Lower the token bill. Keep control of the rollout.
Start with one team, measure the savings, and expand after the results pass your review. Private Mode keeps the original text away from the GPTZIP service.
- Savings tied to measured token reduction
- Client-side masking for sensitive text
- API or proxy rollout by team and workflow
FAQ
What does GPTZIP do?+
GPTZIP shortens eligible prompts before they reach an LLM. A smaller request can use fewer input tokens, cost less, and take less data through the model request.
What does “lossless” mean here?+
GPTZIP checks whether the evaluated task result stays the same after compression. The benchmark section shows the before and after token count beside the checked result. You should still test GPTZIP on your own prompts before a broad rollout.
How fast is GPTZIP?+
In the current 330-token speed comparison, the GPTZIP compression pass takes 30 ms. The stated LLMLingua2 pass takes 1200 ms. Results change with prompt length, mode, hardware, and network setup.
How does GPTZIP lower AI costs?+
LLM providers usually charge for input tokens. When GPTZIP removes eligible input tokens, the billable input can be smaller. Your actual savings depend on the model price, the prompts you send, and the compression measured on your traffic.
How do developers add GPTZIP?+
Use the GPTZIP API before the model call, or point a supported OpenAI-compatible client to GPTZIP Proxy. The API fits apps that own the request code. The proxy fits tools that let you change the base URL.
Does GPTZIP work with LLM harnesses?+
GPTZIP works with most harnesses that support a normal API step or an OpenAI-compatible base URL. Exact support depends on how the harness lets you configure providers and proxies.
What is GPTZIP Private Mode?+
Private Mode runs the masking method in the customer’s browser, computer, or server. It changes selected original terms before the compression request. GPTZIP receives the masked version, while the information needed to restore the terms stays on the customer-controlled device.
Does Private Mode hide data from my LLM provider?+
Private Mode keeps the original text away from the GPTZIP compression service. It does not automatically change what your chosen LLM provider receives later in the workflow. That depends on how your team routes the final request.
Does GPTZIP replace prompt caching or RAG?+
No. Prompt caching lowers the cost of repeated prompt prefixes. RAG selects context. Model routing selects a model. GPTZIP makes eligible prompt text smaller. Teams can use these tools together.
How much setup does GPTZIP need?+
For the API, add one compression call before the LLM request. For GPTZIP Proxy, configure a supported client to use its compatible base URL. Start with a few real prompts, check the result, and then expand.
Next request
Cut the token bill, not the AI work.
Get access to the GPTZIP API, GPTZIP Proxy, and Private Mode.