Compress the tokens

The only
lossless token compression.

It's also the fastest.

Direct API or GPTZIP Proxy Works with most LLM harnesses Private Mode masking

Before GPTZIP

182

tokens

After

129

tokens

kept removed
182 before129 after

29.1%

input reduction

27 ms

Compression Time

Counting GPTZIP token savings…

AI savings calculator

What could fewer tokens save your business?

Pick the work your team does and enter your monthly AI API spend. We'll estimate the impact of sending smaller inputs.

1

Which AI workflow are you paying for?

2

Enter the input-token portion of your usage-based API spend. Do not include seats or subscriptions.

per month

Your potential savings

A smaller input bill, every month.

Enter your monthly spend to see your estimate.

Your result updates instantly.

Projection uses the 19.9% median token reduction found in the closest landing-page benchmark workloads: CoQA · SQuAD v2 · GPQA Diamond · FinanceBench · HumanEval. Output-token charges, subscriptions, seats, and other fees are not reduced.

Benchmark data

LLM accuracy increases after GPTZIP compression.

Accuracy and compression

-20pp
-10pp
0ppNo accuracy drop
+10pp
+20pp
+30pp
+40pp
0%
10%
20%
30%
40%
50%
60%
70%
Median tokens removedAccuracy change vs uncompressed

Methods

GPTZIPHeadroomBear (TheTokenCompany)LLM-Lingua (MSFT)

Datasets

FinanceBenchSQuAD v2CoQAGPQA Diamond
Accuracy and token compression values shown in the chart
BenchmarkMethodMedian tokens removedAccuracy change
FinanceBenchGPTZIP19.4%+6.6pp
FinanceBenchHeadroom16.1%−12.5pp
FinanceBenchBear (TheTokenCompany)1.9%+4.1pp
FinanceBenchLLM-Lingua (MSFT)59.2%−1.6pp
SQuAD v2GPTZIP32.8%+17.2pp
SQuAD v2Headroom10.8%+24.7pp
SQuAD v2Bear (TheTokenCompany)9.3%+34.3pp
SQuAD v2LLM-Lingua (MSFT)61.0%+2.8pp
CoQAGPTZIP42.6%+12.0pp
CoQAHeadroom51.8%−9.7pp
CoQABear (TheTokenCompany)13.1%+16.4pp
CoQALLM-Lingua (MSFT)62.8%−7.6pp
GPQA DiamondGPTZIP19.9%+5.2pp
GPQA DiamondHeadroom10.6%+13.8pp
GPQA DiamondBear (TheTokenCompany)4.9%+37.9pp
GPQA DiamondLLM-Lingua (MSFT)48.4%+4.0pp

Method comparison

GPTZIP is the only lossless compression across use-cases.

Benchmark accuracy and compression comparison for Llama-3.2 4b
BenchmarkUncompressedGPTZIPHeadroomBear (TheTokenCompany)LLM-Lingua (MSFT)
FinanceBench86.7%Accuracy
+6.6ppLossless
93.3% accuracy18.8% compression rate19.4% median40.1% max
−12.5pp
74.2% accuracy16.1% median83.8% max
+4.1ppLossless
90.8% accuracy1.9% median14.1% max
−1.6pp
85.1% accuracy59.2% median68.3% max
LongBench v2N/AAccuracyNot availableNot availableNot availableNot available
SQuAD v256.1%Accuracy
+17.2ppLossless
73.3% accuracy33.5% compression rate32.8% median41.4% max
+24.7ppLossless
80.8% accuracy10.8% median51.1% max
+34.3ppLossless
90.4% accuracy9.3% median23.1% max
+2.8ppLossless
58.9% accuracy61.0% median67.6% max
CoQA76.8%Accuracy
+12.0ppLossless
88.8% accuracy41.8% compression rate42.6% median55.2% max
−9.7pp
67.1% accuracy51.8% median77.3% max
+16.4ppLossless
93.2% accuracy13.1% median25.9% max
−7.6pp
69.2% accuracy62.8% median66.0% max
GPQA Diamond29.3%Accuracy
+5.2ppLossless
34.5% accuracy22.5% compression rate19.9% median76.1% max
+13.8ppLossless
43.1% accuracy10.6% median95.4% max
+37.9ppLossless
67.2% accuracy4.9% median19.8% max
+4.0ppLossless
33.3% accuracy48.4% median78.9% max

Workload fit

Compression depends on the work you send.

Long contexts and short questions do not compress the same way. These rows use the closest measured public benchmark instead of one number for every workload.

Measured GPTZIP workload fit for Llama-3.2 4b
WorkloadPublic proxyMedian removedAccuracy change
Financial and regulatory analysisFinanceBench19.4%+6.6pp
Conversational and extractive QACoQA · SQuAD v232.8–42.6%+12.0pp to +17.2pp
Expert reasoningGPQA Diamond19.9%+5.2pp

Compression speed

330-token measured prompt

GPTZIP
30 ms
Headroom
200 ms
Bear (TheTokenCompany)
250 ms
LLM-Lingua (MSFT)
1200 ms

Terminal-Bench 2.1

Lossless task outcomes with measured cost savings.

GPTZIP keeps the verifier outcome on the shown task pairs while using fewer billed tokens and dollars. The view combines task resolution, cost per successful task, turns, and cache tiers so the compression result is auditable.

GPT-5.6-luna MAXCodex agentReport 2026-08-31

The tasks displayed were chosen at a random draw, more task results are available upon request.

Cost saved

53.7% lower$2.47$4.59 → $2.12 · control → GPTZIP

Tokens saved

67.6% lower6.47M9.57M → 3.10M · control → GPTZIP

Billed turns

57.0% lower85149 → 64 · control → GPTZIP

Quality held

100%3/3 shown pairs3 of 5 paired tasks shown

Outcome economics

Cost per successful task

Control$1.53per shown success
GPTZIP$0.71per shown success

Successful task pairs

Results by task and programming language

Successful Terminal-Bench task results
TaskProgramming languageBilled turnsTotal tokensCostCache read
write-compressorCreate a compressed payload accepted by an existing C decompressor within the size limit.
Any (Rust reference)39 → 1171.8% lower1.81M → 294.7k83.7% lower$1.27 → $0.3373.8% lower1.69M → 257.3k84.8% lower
bn-fit-modifyRecover a Bayesian-network DAG, intervene on Y, and produce the requested sample files.
R-oriented17 → 1041.2% lower288.2k → 168.1k41.7% lower$0.28 → $0.2413.7% lower251.6k → 128.5k48.9% lower
build-cython-extBuild pyknotid Cython extensions and make the package work with the installed NumPy version.
Python/Cython93 → 4353.8% lower7.47M → 2.64M64.7% lower$3.05 → $1.5549.0% lower7.27M → 2.46M66.1% lower

Private prompt compression

Compress private prompts safely.

The masking method runs on your browser, computer, or server. GPTZIP gets the masked text—not the original text. Drag the slider to see the live example.

Original input: Customer support case for Aurora Health: Maya Chen reported a duplicate charge on order 483901. The support agent must verify the billing event, preserve the refund policy, and answer whether the customer can keep the replacement device. Customer support case for Aurora Health: Maya Chen reported a duplicate charge on order 483901. The support agent must verify the billing event, preserve the refund policy, and answer whether the customer can keep the replacement device. Masked text: Client help type in Dawn Wellness: Alice David covered some twin care at club 145302. Any help factor need control each request issue, save each payment insurance, but reply house each buyer may hold each switch twist. Client help type in Dawn Wellness: Alice David covered some twin care at club 145302. Any help factor need control each request issue, save each payment insurance, but reply house each buyer may hold each switch twist.

Drag to compare original and masked text.

What leaves your device

The original data does not go to GPTZIP.

The customer-controlled device keeps the original text and the information needed to restore it. Only masked text goes to the GPTZIP service.

01

Your browser

The original prompt and identifiers stay on your device.

Original text stays here
02

GPTZIP server

Receives only masked text and returns compressed masked output.

Masked text only
03

Compressed output

A shorter masked request comes back to your device.

68 tokens sent
04

Restore in the browser

The browser puts the original terms back into the user-facing text.

Restored locally

88

masked input

68

sent request

20

tokens saved

Your device restores the original terms after the GPTZIP step.
01 / BrowserOriginal inputReadable terms stay local.Customer support case for Aurora Health: Maya Chen reported a duplicate charge on order 483901. The support agent must verify the billing event, preserve the refund policy, and answer whether the customer can keep the replacement device. Customer support case for Aurora Health: Maya Chen reported a duplicate charge on order 483901. The support agent must verify the billing event, preserve the refund policy, and answer whether the customer can keep the replacement device.
Local only

For developers

Less tokens. More LLM quality.

Add GPTZIP before the LLM call or use GPTZIP Proxy. You do not need to change models or rebuild your harness.

  • One API step for apps you own
  • OpenAI-compatible proxy path for supported clients
  • Before-and-after data you can inspect

For enterprise

Lower the token bill. Keep control of the rollout.

Start with one team, measure the savings, and expand after the results pass your review. Private Mode keeps the original text away from the GPTZIP service.

  • Savings tied to measured token reduction
  • Client-side masking for sensitive text
  • API or proxy rollout by team and workflow

FAQ

What does GPTZIP do?+

GPTZIP shortens eligible prompts before they reach an LLM. A smaller request can use fewer input tokens, cost less, and take less data through the model request.

What does “lossless” mean here?+

GPTZIP checks whether the evaluated task result stays the same after compression. The benchmark section shows the before and after token count beside the checked result. You should still test GPTZIP on your own prompts before a broad rollout.

How fast is GPTZIP?+

In the current 330-token speed comparison, the GPTZIP compression pass takes 30 ms. The stated LLMLingua2 pass takes 1200 ms. Results change with prompt length, mode, hardware, and network setup.

How does GPTZIP lower AI costs?+

LLM providers usually charge for input tokens. When GPTZIP removes eligible input tokens, the billable input can be smaller. Your actual savings depend on the model price, the prompts you send, and the compression measured on your traffic.

How do developers add GPTZIP?+

Use the GPTZIP API before the model call, or point a supported OpenAI-compatible client to GPTZIP Proxy. The API fits apps that own the request code. The proxy fits tools that let you change the base URL.

Does GPTZIP work with LLM harnesses?+

GPTZIP works with most harnesses that support a normal API step or an OpenAI-compatible base URL. Exact support depends on how the harness lets you configure providers and proxies.

What is GPTZIP Private Mode?+

Private Mode runs the masking method in the customer’s browser, computer, or server. It changes selected original terms before the compression request. GPTZIP receives the masked version, while the information needed to restore the terms stays on the customer-controlled device.

Does Private Mode hide data from my LLM provider?+

Private Mode keeps the original text away from the GPTZIP compression service. It does not automatically change what your chosen LLM provider receives later in the workflow. That depends on how your team routes the final request.

Does GPTZIP replace prompt caching or RAG?+

No. Prompt caching lowers the cost of repeated prompt prefixes. RAG selects context. Model routing selects a model. GPTZIP makes eligible prompt text smaller. Teams can use these tools together.

How much setup does GPTZIP need?+

For the API, add one compression call before the LLM request. For GPTZIP Proxy, configure a supported client to use its compatible base URL. Start with a few real prompts, check the result, and then expand.

Next request

Cut the token bill, not the AI work.

Get access to the GPTZIP API, GPTZIP Proxy, and Private Mode.