Cohere Rerank vs Jina Reranker vs Pinecone Rerank: Real Reranking Cost per 100 Million Documents in 2026
The short answer: Cohere and Jina meter reranking in two fundamentally different units — Cohere bills per search (a query plus up to 100 documents, or 500-token chunks of them), while Jina bills per token processed. At an ILLUSTRATIVE 200 tokens per document, reranking 100 million documents costs about $2,000 on Cohere Rerank 4 Fast, $2,500 on Rerank 4 Pro, and $1,000 on Jina Reranker v3.5 — Jina coming in half the price of Cohere's cheaper tier at this document length. That ranking does not simply flip once and stay flipped as documents get longer, but it also doesn't alternate forever: against Cohere Fast, the cheaper option alternates within each 500-token chunk band (Jina ahead for the first four-fifths, Fast ahead for the last fifth) only through the first four bands, after which Fast stays permanently cheaper past 2,000 tokens; against Cohere Pro, Jina remains the cheaper option at every document length in this model, with no permanent crossover at all. Pinecone Rerank is not a fourth independent price: it hosts Cohere's models at Cohere's own per-search rate, plus a separate token-billed open-weight model whose exact rate was not published.
What each vendor bills
Cohere (OFFICIAL, Cohere's pricing documentation). Rerank 4 Fast: $2.00 per 1,000 searches. Rerank 4 Pro: $2.50 per 1,000 searches. One search is defined as a single query plus up to 100 documents; any document longer than roughly 500 tokens (including the query) is automatically split into ~500-token chunks, and each chunk counts toward that 100-document limit, so a request with long documents can consume more than one billed search. This makes Cohere's real per-document cost depend on document length in a stepped, not linear, way: documents under 500 tokens cost the same regardless of exact length, but crossing that threshold adds another full chunk-slot.
Jina AI (OFFICIAL, Jina's pricing documentation). Jina Reranker v3.5 (released July 27, 2026, a 0.6-billion-parameter listwise reranker with a 131,072-token context window): $0.05 per million input tokens, output free. This model can rerank up to 64 documents together in a single listwise pass rather than Cohere's pairwise-style scoring, and its huge context window means very long documents do not need chunking the way Cohere's do. An older model, Jina Reranker v2 Base Multilingual, remains available at $0.018 per million tokens for both input and output (a 1,024-token context window, materially smaller than v3.5's), and open weights for Jina's rerankers are separately available under a non-commercial license for self-hosting.
Pinecone Rerank (OFFICIAL for the mechanism, REPORTED for pricing). Pinecone's Inference API hosts third-party and open-weight rerank models rather than running its own proprietary reranker. It offers cohere-rerank-4-fast as a selectable model, billed at Cohere's own $2.00-per-1,000-search rate (Pinecone passes this through rather than repricing it), and a separate open-weight model, bge-reranker-v2-m3, billed by tokens used, whose exact per-token rate was not published in the documentation reviewed (QUOTE-ONLY / UNKNOWN). Rate limits on Pinecone's free Starter plan cap bge-reranker-v2-m3 usage at 5 million tokens; paid-tier throughput requires a support request for a rate increase.
Normalizing the unit: documents, searches, and tokens
All ILLUSTRATIVE. This article uses 200 tokens per document as a baseline chunk length (a typical paragraph-to-page-sized retrieval chunk) and 100 documents per Cohere search (Cohere's own stated batch limit). Formula: Cohere searches = documents ÷ 100; Jina tokens = documents × 200.
Cost per volume
Formula: Cohere cost = (searches ÷ 1,000) × rate per 1,000 searches; Jina cost = (tokens ÷ 1,000,000) × $0.05.
| Documents reranked | Cohere Rerank 4 Fast | Cohere Rerank 4 Pro | Jina Reranker v3.5 |
|---|---|---|---|
| 1,000,000 | $20.00 | $25.00 | $10.00 |
| 10,000,000 | $200.00 | $250.00 | $100.00 |
| 100,000,000 | $2,000.00 | $2,500.00 | $1,000.00 |
At this 200-token document length, Jina Reranker v3.5 is consistently half the cost of Cohere's cheaper Fast tier and 40% of the Pro tier's cost, at every volume, because 200 tokens sits well within the cheaper first four-fifths of Cohere's first 500-token chunk band (see the sensitivity section below for how this changes at other document lengths).
Where the ranking flips: document length
Cohere's $2.00 per 1,000 searches applies to a search (up to 100 chunk-slots), not to one document directly. Formula: Cohere cost per chunk-slot = ($2.00 ÷ 1,000 searches) ÷ 100 chunk-slots per search = $0.00002, and a document consumes one chunk-slot per ~500 tokens, rounded up (a 350-token document costs the same $0.00002 as a 10-token document; a 600-token document needs 2 chunk-slots and costs $0.00004). Jina, by contrast, bills smoothly per token: $0.05 ÷ 1,000,000 = $0.00000005 per token.
This produces a stepped comparison that only alternates for a few chunk bands, not indefinitely. Formula: break-even tokens for Cohere Fast at chunk band n = n × 400 (Cohere's flat per-n-chunk cost, $0.00002n, is worth exactly 400n tokens of Jina's linear rate). Chunk band n covers tokens from 500(n−1)+1 to 500n, so the break-even falls inside the band only while 400n is less than the band's own end (500n) — true for n = 1 through 4, where the crossover lands at 400, 800, 1,200 and 1,600 tokens respectively. From band 5 onward (documents of 2,001 tokens or more), the arithmetic flips permanently: at 2,001 tokens Jina already costs $0.00010005 against Cohere Fast's flat $0.0001 for that 5-chunk band, so Cohere Fast is the cheaper option throughout band 5 and every band after it, with no further alternation back to Jina.
| Document length | Chunk-slots (Cohere) | Cohere Fast cost/document | Jina cost/document | Cheaper option |
|---|---|---|---|---|
| 1–400 tokens | 1 | $0.00002 | up to $0.00002 | Jina |
| 400–500 tokens | 1 | $0.00002 | $0.00002–$0.000025 | Cohere Fast |
| 501–800 tokens | 2 | $0.00004 | $0.000025–$0.00004 | Jina |
| 800–1,000 tokens | 2 | $0.00004 | $0.00004–$0.00005 | Cohere Fast |
| 1,001–1,200 tokens | 3 | $0.00006 | $0.00005–$0.00006 | Jina |
| 1,200–1,500 tokens | 3 | $0.00006 | $0.00006–$0.000075 | Cohere Fast |
| 1,501–1,600 tokens | 4 | $0.00008 | $0.000075–$0.00008 | Jina |
| 1,600–2,000 tokens | 4 | $0.00008 | $0.00008–$0.0001 | Cohere Fast |
| 2,001 tokens and longer | 5 or more | $0.00002 per chunk, rising | continues linearly, now always ahead of Cohere at every band's start | Cohere Fast, permanently |
Cohere Pro behaves differently because its per-chunk-slot cost is higher: $2.50 ÷ 1,000 searches ÷ 100 chunk-slots = $0.000025, exactly 500 tokens' worth of Jina's rate rather than 400. That means the break-even for Pro at band n is 500n tokens — precisely the band's own upper boundary — so Jina is cheaper for every token strictly inside a band and the two only tie exactly at each 500-token multiple (500, 1,000, 1,500, and so on). Unlike Fast, Pro never crosses over into being the cheaper option for any sustained range of document lengths; Jina remains at or below Pro's cost at every document length in this model.
At the article's 200-token baseline, Jina is cheaper than both Cohere tiers ($0.00001 versus Cohere Fast's $0.00002), which is why the 100-million-document table above shows Jina at roughly half of Cohere Fast. Whether Jina stays cheaper as documents get longer depends on which Cohere tier is the comparison: against Fast, Jina loses ground inside the last fifth of chunk bands 1 through 4 and loses permanently past 2,000 tokens; against Pro, Jina remains the cheaper (or tied) option at every document length modeled here, with no permanent crossover at all.
Listwise versus pairwise reranking
Jina Reranker v3.5's listwise architecture scores up to 64 candidate documents together in one pass, which independent benchmarking associates with better handling of relative ranking among similar candidates than a pairwise cross-encoder scoring each document independently against the query. Cohere's Rerank 4 models use a more conventional per-document scoring approach. Neither architecture choice is priced directly — it shows up in relevance quality (reported 15–30% RAG answer-quality improvement from adding any reranker over vector search alone), not in the dollar figures above.
Sensitivity
- Document length. The single largest lever for Jina's cost; a corpus of 1,000-token documents costs Jina 5 times what a 200-token corpus costs, while Cohere's cost is far less sensitive to length until the 500-token chunk boundary.
- Candidates per query. Reranking 200 candidates instead of 100 per query doubles Cohere's search count (two chunks of documents = two searches) and doubles Jina's token count identically.
- Model tier. Cohere Pro costs 25% more than Fast for (per Cohere's own positioning) better accuracy on harder ranking tasks; Jina's v2 and v3.5 models differ by context window and architecture more than by price.
- Pinecone's unpriced bge model. Teams already inside Pinecone's ecosystem may default to bge-reranker-v2-m3 for convenience without knowing its per-token rate, which was not published.
Budgeting traps
- Assuming reranking cost scales with query volume alone. It scales with the number of candidate documents reranked per query, which is a separate and often larger number.
- Comparing a per-search price to a per-token price without normalizing document length. The two billing mechanisms are not directly comparable without an explicit tokens-per-document assumption.
- Reranking too many candidates. Reranking thousands of candidates per query is both expensive and, per practitioner guidance, rarely improves quality versus narrowing the first-stage retrieval set to 50–200 candidates before reranking.
- Treating Pinecone Rerank as an independently priced product. Its Cohere-model option bills at Cohere's own rate; its open-weight option has an unpublished rate.
What to ask before you buy
Measure your actual average document or chunk length before comparing Cohere's per-search price to Jina's per-token price, since the crossover depends entirely on that number. Ask Pinecone directly for the current per-token rate on bge-reranker-v2-m3 if you plan to rely on that specific model rather than its Cohere pass-through option.