How the benchmark works

Short sentences carrying a number for a city — the needles — are planted inside long Turkish Wikipedia text — the haystack. The model then has to report every city's number. Below is one complete, real test prompt: the smallest cell of the grid (1,000 tokens, two needles), rebuilt with the benchmark code, and gemma-4-31B-it's actual answer.

1The haystack, with two needles planted
== Aldebaran == Aldebaran (α Tau, α Tauri, Alpha Tauri), Boğa takımyıldızı yönünde yaklaşık olarak 65 ışık yılı uzaklıkta bulunan bir kırmızı dev yıldızdır. 0,87 […] […] Özellikler. K5III tipi bir yıldız olarak sınıflandırılan Aldebaran, Hertzsprung-Russell diyagramı'nın Anakol çizgisinden uzaklaşmış bir kırmızı dev yıldızdır. Çekirdeğindeki hidrojen yakıtı tükenmiş ve hidrojen füzyonu durmuştur. Yangon şehrinin özel sihirli sayısı: 176needle 1 · 32% depth helyum füzyonu için yeterince sıcak olmamasına rağmen, yıldızın çekirdek ısısı kütleçekimi basıncı nedeniyle artmış ve çapı Güneş'in çapının 44,2 katı artmıştır (yaklaşık 61 milyon kilometre). Hipparcos uydusu, uzaklığını 65,1 ışık yılı […] […] O'nu; Betelgeuse (-2,9), R Doradus (-2,6) ve Arcturus'dan (-2,2) daha parlak yapmaktadır. Gözlem bilgisi. Sofia şehrinin özel sihirli sayısı: 260needle 2 · 63% depth Aldebaran, kısmen parlaklığı ve kısmen de çok belirgin bir yıldız deseniyle uzaysal ilişkisi nedeniyle, gece gökyüzünde bulması kolay olan yıldızlardan biridir. Avcı'nın Kuşağı'nı oluşturan üç […]
Show the full 799-token text
== Aldebaran == Aldebaran (α Tau, α Tauri, Alpha Tauri), Boğa takımyıldızı yönünde yaklaşık olarak 65 ışık yılı uzaklıkta bulunan bir kırmızı dev yıldızdır. 0,87 kadir ortalama büyüklüğüyle takımyıldızının en parlak yıldızı olan Aldebaran, gece gökyüzünün de en parlak yıldızlarından birisidir. Aldebaran, Arapça (الدبران) 'takip eden' veya 'izleyen' demektir. Bu konumda 'izlenen' "V" harfi şeklinde bir araya gelmiş bir yıldız kümesi olan Hyades'dir. 1997 yılında bir altyıldız arkadaşı olduğu iddia edilmiş, fakat sonradan yapılan gözlemler bu iddiayı doğrulamamıştır. Fiziksel Özellikler. K5III tipi bir yıldız olarak sınıflandırılan Aldebaran, Hertzsprung-Russell diyagramı'nın Anakol çizgisinden uzaklaşmış bir kırmızı dev yıldızdır. Çekirdeğindeki hidrojen yakıtı tükenmiş ve hidrojen füzyonu durmuştur. Yangon şehrinin özel sihirli sayısı: 176needle 1 · 32% depth helyum füzyonu için yeterince sıcak olmamasına rağmen, yıldızın çekirdek ısısı kütleçekimi basıncı nedeniyle artmış ve çapı Güneş'in çapının 44,2 katı artmıştır (yaklaşık 61 milyon kilometre). Hipparcos uydusu, uzaklığını 65,1 ışık yılı (20.0 pc) olarak ölçmüştür ve Güneş'ten 425 kat fazla aydınlatma gücüne sahiptir. Aldebaran, "LB" tipi yavaş düzensiz değişen bir yıldızdır. Görünen büyüklüğü 0,2'ye göre yaklaşık olarak 0,75 ile 0,95 aralığında değişir. Yakın-kızılötesi J bandı büyüklüğü -2,1'dir ve bu büyüklük O'nu; Betelgeuse (-2,9), R Doradus (-2,6) ve Arcturus'dan (-2,2) daha parlak yapmaktadır. Gözlem bilgisi. Sofia şehrinin özel sihirli sayısı: 260needle 2 · 63% depth Aldebaran, kısmen parlaklığı ve kısmen de çok belirgin bir yıldız deseniyle uzaysal ilişkisi nedeniyle, gece gökyüzünde bulması kolay olan yıldızlardan biridir. Avcı'nın Kuşağı'nı oluşturan üç yıldızdan, soldan sağa (kuzey yarımkürede) doğru bir hat izlendiğinde bulunan en parlak yıldız Aldebaran'dır. Aldebaran, boğa'nın başını oluşturan V şeklindeki açık yıldız kümesi Hyades'in en parlak üyesidir. Ancak Aldebaran, sadece şans eseri Dünya ile Hyades arasındaki görüş doğrultusunda yer alır. Yıldız kümesi (Hyades), Dünya'dan 150 ışık yılı uzaklığıyla Aldebaran'a göre iki kat daha fazla uzakta yer alır. Aldebaran, Ay tarafından örtülmesine yeterli olacak kadar tutulum düzlemine yakındır. Ay'ın yükselen düğümü, sonbahar ekinoksuna yakın olduğunda böyle bir örtülme gerçekleşebilir. Böyle bir olay, 2015 yılı civarında meydana gelecektir. Aldebaran'ın çapının daha doğru bir şekilde tahmin edilmesi 22 Eylül 1978 tarihinde gerçekleşen bir örtülmeyle olmuştur.

Article: Aldebaran (799 tokens). Each needle reads “The special magic number of the city of Yangon: 176”. Needles are spread evenly through the text and moved back to the nearest sentence end.

2The question
Aşağıdaki şehirlerin özel sihirli sayıları nedir: Yangon, Sofia? Her şehir için bulduğun sayıyı listele.

“What are the special magic numbers of the following cities: Yangon, Sofia? List the number you found for each city.” The prompt also fixes the answer format: one SEHIR: <city> = <number> line per city, YOK if not found.

3The model's answer
SEHIR: Yangon = 176
SEHIR: Sofia = 260

gemma-4-31B-it, verbatim → 2 / 2 needles retrieved.

4How an answer line is scored
  • ✓Yangon = 176retrieved — right number, right city
  • ✕Yangon = YOKmissed — reported as not found
  • ✕Yangon = 260missed — a real needle, but Sofia's number

The last case is why scoring is city-bound: Gemma-3 models often find the right numbers but attach them to the wrong cities.

This is repeated at 15 context lengths (1,000 → 127,000 tokens, and up to 261,000) and for many needle counts. Each heatmap cell below is one prompt like this one, coloured by the share of needles retrieved.

Heatmaps

Each cell is one prompt like the example above — one context length × one condition — coloured by the share of needles retrieved. Hover any cell for the exact result.

N needles, each a Turkish sentence of the form “<city> şehrinin özel sihirli sayısı: <number>”, are spread evenly through the haystack; the model must return every city's number. A needle counts only when its number is written against its own city. Grid: 15 context lengths × N ∈ {1, 2, 3, 5, 7, 10}. Ordered by accuracy. The 5 Qwen models were added later on the identical protocol (same key pool, value width and grid), answering with thinking disabled.

gemma-4-31B-it
420 / 420 needles · 100.0%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
2100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
3100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
5100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
7100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
10100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
context length (thousand tokens) →
Qwen3.8-27B
420 / 420 needles · 100.0%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
2100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
3100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
5100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
7100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
10100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
context length (thousand tokens) →
Qwen3.6-35B-A3B
419 / 420 needles · 99.8%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
2100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
3100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
5100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%8080.0%100.0%100.0%
7100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
10100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
context length (thousand tokens) →
Qwen3.5-9B
411 / 420 needles · 97.9%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%00.0%100.0%
2100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
3100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%6766.7%100.0%100.0%
5100.0%100.0%100.0%100.0%100.0%8080.0%100.0%8080.0%100.0%100.0%100.0%100.0%8080.0%100.0%100.0%
7100.0%100.0%100.0%100.0%100.0%100.0%8685.7%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
10100.0%100.0%100.0%100.0%100.0%100.0%9090.0%100.0%9090.0%100.0%100.0%9090.0%100.0%100.0%100.0%
context length (thousand tokens) →
gemma-4-E2B-it
408 / 420 needles · 97.1%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%00.0%100.0%100.0%100.0%
2100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
3100.0%100.0%100.0%100.0%100.0%100.0%100.0%6766.7%100.0%100.0%100.0%100.0%6766.7%100.0%100.0%
5100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
7100.0%100.0%100.0%100.0%100.0%100.0%7171.4%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
10100.0%100.0%100.0%100.0%100.0%100.0%9090.0%100.0%100.0%100.0%9090.0%8080.0%7070.0%100.0%100.0%
context length (thousand tokens) →
Qwen3.5-2B
407 / 420 needles · 96.9%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
2100.0%100.0%100.0%100.0%5050.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
3100.0%100.0%100.0%100.0%3333.3%100.0%100.0%100.0%100.0%3333.3%100.0%6766.7%100.0%100.0%100.0%
5100.0%8080.0%100.0%100.0%100.0%100.0%8080.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%8080.0%
7100.0%100.0%100.0%8685.7%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%8685.7%100.0%8685.7%
10100.0%100.0%100.0%100.0%100.0%100.0%9090.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
context length (thousand tokens) →
Qwen3.5-4B
400 / 420 needles · 95.2%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%00.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
2100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
3100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%6766.7%100.0%3333.3%
5100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
7100.0%100.0%100.0%7171.4%7171.4%100.0%8685.7%100.0%100.0%100.0%100.0%100.0%100.0%8685.7%8685.7%
10100.0%100.0%100.0%8080.0%100.0%100.0%9090.0%9090.0%100.0%9090.0%9090.0%9090.0%9090.0%9090.0%100.0%
context length (thousand tokens) →
gemma-4-12B-it
388 / 420 needles · 92.4%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
2100.0%100.0%100.0%100.0%100.0%5050.0%100.0%100.0%100.0%5050.0%100.0%100.0%100.0%5050.0%5050.0%
3100.0%100.0%6766.7%100.0%6766.7%100.0%100.0%100.0%100.0%100.0%6766.7%100.0%100.0%6766.7%100.0%
5100.0%100.0%100.0%100.0%8080.0%100.0%100.0%100.0%100.0%100.0%100.0%8080.0%100.0%8080.0%100.0%
7100.0%100.0%100.0%8685.7%8685.7%8685.7%100.0%8685.7%8685.7%100.0%8685.7%100.0%7171.4%7171.4%8685.7%
10100.0%100.0%100.0%100.0%100.0%100.0%9090.0%9090.0%9090.0%8080.0%8080.0%9090.0%9090.0%100.0%9090.0%
context length (thousand tokens) →
gemma-4-26B-A4B-it
378 / 420 needles · 90.0%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%00.0%100.0%100.0%100.0%
2100.0%100.0%100.0%100.0%5050.0%100.0%100.0%100.0%100.0%100.0%100.0%5050.0%5050.0%100.0%100.0%
3100.0%100.0%100.0%100.0%100.0%100.0%6766.7%100.0%6766.7%100.0%6766.7%3333.3%6766.7%100.0%6766.7%
5100.0%100.0%100.0%100.0%8080.0%8080.0%8080.0%100.0%100.0%100.0%100.0%8080.0%100.0%8080.0%100.0%
7100.0%100.0%8685.7%100.0%8685.7%100.0%7171.4%8685.7%7171.4%8685.7%8685.7%7171.4%100.0%8685.7%8685.7%
10100.0%100.0%9090.0%100.0%9090.0%9090.0%9090.0%8080.0%100.0%8080.0%100.0%9090.0%9090.0%8080.0%9090.0%
context length (thousand tokens) →
gemma-4-E4B-it
351 / 420 needles · 83.6%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%00.0%00.0%100.0%00.0%00.0%100.0%
2100.0%100.0%100.0%100.0%100.0%5050.0%100.0%5050.0%5050.0%5050.0%5050.0%5050.0%5050.0%5050.0%5050.0%
3100.0%100.0%100.0%6766.7%100.0%100.0%100.0%6766.7%6766.7%3333.3%6766.7%6766.7%3333.3%6766.7%3333.3%
5100.0%100.0%100.0%8080.0%8080.0%8080.0%8080.0%6060.0%100.0%8080.0%8080.0%8080.0%8080.0%8080.0%8080.0%
7100.0%100.0%8685.7%100.0%100.0%8685.7%8685.7%8685.7%8685.7%7171.4%7171.4%7171.4%7171.4%7171.4%7171.4%
10100.0%100.0%100.0%100.0%100.0%9090.0%9090.0%9090.0%100.0%7070.0%100.0%9090.0%8080.0%7070.0%7070.0%
context length (thousand tokens) →
gemma-3-27b-it
235 / 420 needles · 56.0%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%00.0%100.0%100.0%00.0%00.0%00.0%00.0%00.0%00.0%
2100.0%100.0%100.0%100.0%100.0%100.0%5050.0%00.0%5050.0%00.0%5050.0%00.0%5050.0%00.0%5050.0%
3100.0%100.0%100.0%100.0%100.0%6766.7%3333.3%100.0%6766.7%3333.3%3333.3%3333.3%6766.7%3333.3%3333.3%
5100.0%8080.0%100.0%100.0%4040.0%2020.0%2020.0%4040.0%00.0%6060.0%6060.0%4040.0%2020.0%6060.0%00.0%
7100.0%8685.7%8685.7%7171.4%5757.1%7171.4%5757.1%5757.1%4342.9%1414.3%1414.3%7171.4%4342.9%4342.9%5757.1%
109090.0%100.0%8080.0%100.0%8080.0%5050.0%6060.0%3030.0%4040.0%4040.0%4040.0%2020.0%1010.0%2020.0%6060.0%
context length (thousand tokens) →
gemma-3-12b-it
173 / 420 needles · 41.2%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%00.0%00.0%100.0%100.0%00.0%100.0%00.0%
2100.0%5050.0%5050.0%100.0%100.0%5050.0%100.0%100.0%5050.0%100.0%100.0%5050.0%5050.0%5050.0%100.0%
3100.0%100.0%100.0%3333.3%100.0%3333.3%3333.3%00.0%6766.7%00.0%6766.7%3333.3%3333.3%100.0%100.0%
5100.0%100.0%100.0%4040.0%00.0%00.0%00.0%4040.0%4040.0%2020.0%4040.0%2020.0%4040.0%6060.0%00.0%
7100.0%100.0%4342.9%4342.9%4342.9%1414.3%5757.1%1414.3%00.0%1414.3%00.0%2928.6%1414.3%1414.3%2928.6%
109090.0%9090.0%5050.0%6060.0%4040.0%00.0%1010.0%00.0%4040.0%00.0%3030.0%2020.0%3030.0%00.0%00.0%
context length (thousand tokens) →
gemma-3-4b-it
150 / 420 needles · 35.7%
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%00.0%100.0%00.0%00.0%00.0%00.0%00.0%00.0%00.0%
2100.0%5050.0%5050.0%00.0%5050.0%5050.0%00.0%00.0%5050.0%5050.0%00.0%00.0%00.0%5050.0%00.0%
3100.0%100.0%100.0%00.0%00.0%00.0%3333.3%3333.3%6766.7%6766.7%3333.3%00.0%00.0%00.0%6766.7%
5100.0%2020.0%2020.0%100.0%4040.0%4040.0%4040.0%2020.0%4040.0%6060.0%2020.0%2020.0%00.0%2020.0%00.0%
78685.7%7171.4%00.0%1414.3%00.0%1414.3%2928.6%1414.3%4342.9%4342.9%2928.6%00.0%2928.6%2928.6%5757.1%
10100.0%6060.0%5050.0%4040.0%5050.0%5050.0%4040.0%2020.0%4040.0%3030.0%2020.0%1010.0%5050.0%1010.0%00.0%
context length (thousand tokens) →
RankModelFamilyNeedles boundAccuracyValue-onlyBinding gapWrong cityMis-copiedSpurious
1gemma-4-31B-itGemma420 / 420100.0%99.8%-1000
1Qwen3.8-27BQwen420 / 420100.0%100.0%+0000
3Qwen3.6-35B-A3BQwen419 / 42099.8%99.8%+0000
4Qwen3.5-9BQwen411 / 42097.9%97.9%+0000
5gemma-4-E2B-itGemma408 / 42097.1%97.1%+0201
6Qwen3.5-2BQwen407 / 42096.9%97.9%+4221
7Qwen3.5-4BQwen400 / 42095.2%97.6%+10190
8gemma-4-12B-itGemma388 / 42092.4%92.4%+0100
9gemma-4-26B-A4B-itGemma378 / 42090.0%90.0%+0000
10gemma-4-E4B-itGemma351 / 42083.6%83.3%-1100
11gemma-3-27b-itGemma235 / 42056.0%66.4%+443716
12gemma-3-12b-itGemma173 / 42041.2%64.0%+9697216
13gemma-3-4b-itGemma150 / 42035.7%61.0%+10696440
Value-only credits a number found anywhere in the answer. The binding gap is how many needles value-only scoring credits that city-bound scoring does not, split into a number written against another requested city and one written against a mis-copied name for its own city. Spurious counts numbers reported that were never planted. Every Gemma-3 model finds the right numbers but attaches them to the wrong cities, which value-only scoring cannot see. The small Qwen gaps are mostly the other kind: Qwen3.5-4B writes Turkish names (BEYRUT, TAŞKENT, BAĞDAD) or typos (HELINSKI), which the scorer does not credit for any model.

How many facts can a model retrieve in one pass? gemma-4-31B-it holds ~99.9% all the way to 419 needles — 419 city→number bindings in a single ~9,300-token answer. Accuracy does not fall with N; the 419-city key pool, not the model, is the limit.

gemma-4-31B-it · N = 1 → 419
21,227 / 21,252 needles · 99.9%
every N from 1 to 419 on the same grid, corpus and weights
N =110192837465564738291100109118127
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
2100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
3100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
5100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
7100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
10100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
15100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
20100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
30100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
40not run100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
60not run100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%9898.3%
90not run100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%9998.9%100.0%9998.9%9998.9%100.0%
130not run100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%9999.2%100.0%100.0%
180not run100.0%100.0%100.0%100.0%100.0%9999.4%100.0%100.0%9998.9%100.0%100.0%100.0%100.0%100.0%
250not runnot run100.0%100.0%100.0%100.0%100.0%99.699.6%100.0%99.699.6%100.0%99.699.6%100.0%99.699.6%99.699.6%
320not runnot run100.0%100.0%100.0%100.0%100.0%100.0%99.699.7%100.0%99.699.7%99.699.7%99.699.7%100.0%100.0%
419not runnot run100.0%100.0%100.0%100.0%99.799.8%100.0%100.0%99.799.8%100.0%99.799.8%99.799.8%99.799.8%9999.3%
context length (thousand tokens) →
20 of the 25 misses are the single key Kabul — the everyday Turkish word for “acceptance”, which occurs 86 times in the haystack. See Findings.
gemma-4-12B-it
4,148 / 4,340 needles · 95.6%
N =10192837465564738291100109118127
40100.0%100.0%100.0%100.0%9897.5%9897.5%8887.5%9897.5%9292.5%8080.0%8282.5%9595.0%9292.5%7575.0%
90100.0%100.0%100.0%9897.8%100.0%100.0%9897.8%9998.9%9998.9%9292.2%8988.9%9090.0%8181.1%7171.1%
180100.0%100.0%100.0%100.0%9998.9%9898.3%9999.4%9998.9%9797.2%9897.8%9392.8%9695.6%9191.1%8887.8%
context length (thousand tokens) →
gemma-4-26B-A4B-it
3,550 / 4,340 needles · 81.8%
N =10192837465564738291100109118127
40100.0%100.0%9595.0%9595.0%7575.0%8887.5%8887.5%7272.5%8887.5%8585.0%7070.0%7877.5%7877.5%5555.0%
90100.0%9897.8%100.0%9796.7%9897.8%9796.7%8685.6%9191.1%8685.6%8383.3%6867.8%5252.2%5857.8%5958.9%
180100.0%9999.4%9695.6%9696.1%9191.1%9090.0%8382.8%8887.8%7676.1%7675.6%6565.0%6060.0%6161.1%4747.2%
context length (thousand tokens) →
ModelNNeedles boundAccuracy
gemma-4-31B-it115 / 15100.00%
gemma-4-31B-it230 / 30100.00%
gemma-4-31B-it345 / 45100.00%
gemma-4-31B-it575 / 75100.00%
gemma-4-31B-it7105 / 105100.00%
gemma-4-31B-it10150 / 150100.00%
gemma-4-31B-it15225 / 225100.00%
gemma-4-31B-it20300 / 300100.00%
gemma-4-31B-it30450 / 450100.00%
gemma-4-31B-it40560 / 560100.00%
gemma-4-31B-it60839 / 84099.88%
gemma-4-31B-it901,257 / 1,26099.76%
gemma-4-31B-it1301,819 / 1,82099.95%
gemma-4-31B-it1802,517 / 2,52099.88%
gemma-4-31B-it2503,245 / 3,25099.85%
gemma-4-31B-it3204,156 / 4,16099.90%
gemma-4-31B-it4195,439 / 5,44799.85%
gemma-4-12B-it40519 / 56092.68%
gemma-4-12B-it901,185 / 1,26094.05%
gemma-4-12B-it1802,444 / 2,52096.98%
gemma-4-26B-A4B-it40466 / 56083.21%
gemma-4-26B-A4B-it901,054 / 1,26083.65%
gemma-4-26B-A4B-it1802,030 / 2,52080.56%

Needle spacing, not needle count, predicts failure. Raising N at fixed context packs needles closer together, so the standard grid cannot separate the two. Here N and context are fixed and only the spread changes (mean needle depth pinned at 50%). Tighter spreads are easier for every model — and fewer needles is harder, the opposite of the usual assumption.

gemma-4-12B-it · N = 10
525 / 560 needles · 93.8%
rows: share of the document the needles are spread over
span10192837465564738291100109118127
100%100.0%100.0%9090.0%100.0%9090.0%9090.0%9090.0%100.0%9090.0%7070.0%9090.0%9090.0%9090.0%7070.0%
50%100.0%100.0%100.0%9090.0%100.0%100.0%9090.0%100.0%100.0%9090.0%9090.0%9090.0%9090.0%9090.0%
25%100.0%100.0%100.0%100.0%100.0%100.0%100.0%9090.0%100.0%8080.0%9090.0%8080.0%7070.0%9090.0%
10%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%8080.0%9090.0%9090.0%
context length (thousand tokens) →
gemma-4-12B-it · N = 20
1,065 / 1,120 needles · 95.1%
rows: share of the document the needles are spread over
span10192837465564738291100109118127
100%100.0%100.0%100.0%100.0%9595.0%9090.0%9595.0%9090.0%8080.0%9090.0%8585.0%9595.0%8585.0%8585.0%
50%100.0%100.0%100.0%100.0%100.0%100.0%9595.0%9090.0%9090.0%9595.0%9595.0%9090.0%9595.0%8080.0%
25%100.0%100.0%100.0%100.0%9595.0%9090.0%100.0%9090.0%9595.0%9090.0%9595.0%8585.0%100.0%9595.0%
10%100.0%100.0%100.0%100.0%100.0%100.0%100.0%9595.0%9595.0%100.0%100.0%100.0%9595.0%8585.0%
context length (thousand tokens) →
gemma-4-12B-it · N = 40
2,159 / 2,240 needles · 96.4%
rows: share of the document the needles are spread over
span10192837465564738291100109118127
100%100.0%100.0%100.0%100.0%9897.5%9897.5%100.0%9292.5%100.0%8282.5%9595.0%8282.5%8887.5%8887.5%
50%100.0%100.0%9897.5%100.0%100.0%9897.5%9897.5%100.0%9090.0%8887.5%9595.0%9897.5%8282.5%7877.5%
25%9897.5%100.0%100.0%100.0%100.0%100.0%100.0%9897.5%100.0%9292.5%9292.5%9090.0%100.0%9897.5%
10%100.0%100.0%100.0%9897.5%100.0%9897.5%100.0%100.0%100.0%100.0%9595.0%9595.0%100.0%100.0%
context length (thousand tokens) →
gemma-4-12B-it · N = 90
4,916 / 5,040 needles · 97.5%
rows: share of the document the needles are spread over
span10192837465564738291100109118127
100%100.0%100.0%100.0%100.0%100.0%9998.9%9897.8%9897.8%100.0%9494.4%9090.0%8685.6%8786.7%8282.2%
50%100.0%100.0%9998.9%9998.9%9897.8%9998.9%9998.9%100.0%9695.6%9494.4%9897.8%9796.7%9494.4%8988.9%
25%100.0%100.0%100.0%100.0%9796.7%100.0%100.0%100.0%9998.9%9998.9%9897.8%9494.4%9796.7%9393.3%
10%100.0%100.0%100.0%100.0%100.0%9897.8%100.0%9998.9%100.0%9998.9%100.0%100.0%9998.9%9796.7%
context length (thousand tokens) →
gemma-4-26B-A4B-it · N = 40
2,025 / 2,240 needles · 90.4%
rows: share of the document the needles are spread over
span10192837465564738291100109118127
100%100.0%100.0%9595.0%9595.0%9090.0%9090.0%9292.5%8585.0%7877.5%8585.0%6262.5%7070.0%6565.0%6060.0%
50%100.0%100.0%100.0%9897.5%100.0%100.0%9595.0%8887.5%8887.5%8080.0%7575.0%7070.0%7272.5%7070.0%
25%100.0%100.0%100.0%100.0%9897.5%9897.5%100.0%9897.5%9090.0%8887.5%9897.5%8887.5%9090.0%8080.0%
10%100.0%100.0%6060.0%100.0%100.0%100.0%9897.5%100.0%9897.5%100.0%9595.0%9897.5%9595.0%9292.5%
context length (thousand tokens) →
gemma-3-12b-it · N = 40
720 / 2,240 needles · 32.1%
rows: share of the document the needles are spread over
span10192837465564738291100109118127
100%9595.0%6060.0%7877.5%4040.0%2525.0%2525.0%1212.5%1515.0%1212.5%00.0%00.0%00.0%1010.0%00.0%
50%9292.5%7070.0%6867.5%2525.0%1212.5%00.0%1817.5%55.0%87.5%00.0%22.5%1010.0%1817.5%00.0%
25%100.0%100.0%8585.0%3030.0%4242.5%1817.5%2827.5%3030.0%1010.0%00.0%00.0%22.5%87.5%00.0%
10%100.0%100.0%8282.5%8080.0%6565.0%7575.0%00.0%4545.0%1212.5%1212.5%2222.5%2020.0%55.0%2827.5%
context length (thousand tokens) →
Needle spacing (tokens)CellsMissedMiss rateNeedle counts in bin
0–2566122 / 3,7100.59%10, 20, 40, 90
256–5123833 / 1,7701.86%10, 20, 40, 90
512–1,0244259 / 1,6103.66%10, 20, 40, 90
1,024–2,0483695 / 1,1008.64%10, 20, 40, 90
2,048–4,0962756 / 52010.77%10, 20, 40
4,096+2030 / 25012.00%10, 20
gemma-4-12B-it, all four needle counts pooled (224 cells, 8,960 needles). Below a 1,024-token gap 1.61% of needles are missed; at or above it 9.68% — about 6×, at the scale of the model's 1,024-token sliding-window attention. Spearman: miss vs spacing ρ = +0.586 (p < 1e-8); miss vs needle count ρ = −0.109 (p = 0.10).

The same retrieval task pushed to 261,000 tokens, the edge of the models' 262,144-token window. gemma-4-31B-it only fits at this length on an 80 GB A100 with fp8 weight quantization, so its numbers are not precision-comparable to the bf16 rows.

gemma-4-31B-it
239 / 240 needles · 99.6%
fp8 weights
N =11938567593112131149168186205223242261
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
5100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%8080.0%100.0%100.0%
10100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%
context length (thousand tokens) →
gemma-4-26B-A4B-it
196 / 240 needles · 81.7%
bf16
N =11938567593112131149168186205223242261
1100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%00.0%100.0%100.0%00.0%00.0%
5100.0%100.0%100.0%8080.0%100.0%100.0%6060.0%100.0%8080.0%100.0%4040.0%100.0%8080.0%8080.0%2020.0%
10100.0%100.0%100.0%9090.0%9090.0%8080.0%8080.0%7070.0%8080.0%6060.0%8080.0%7070.0%8080.0%8080.0%6060.0%
context length (thousand tokens) →
gemma-4-12B-it
186 / 240 needles · 77.5%
bf16
N =11938567593112131149168186205223242261
1100.0%100.0%100.0%100.0%100.0%00.0%100.0%100.0%100.0%00.0%00.0%00.0%00.0%00.0%00.0%
5100.0%100.0%100.0%8080.0%8080.0%8080.0%100.0%8080.0%8080.0%4040.0%2020.0%8080.0%8080.0%8080.0%6060.0%
10100.0%100.0%100.0%9090.0%9090.0%9090.0%100.0%8080.0%100.0%7070.0%7070.0%7070.0%8080.0%00.0%6060.0%
context length (thousand tokens) →
ModelPrecisionNeedles boundAccuracy
gemma-4-31B-itfp8 weights239 / 24099.6%
gemma-4-26B-A4B-itbf16196 / 24081.7%
gemma-4-12B-itbf16186 / 24077.5%

Findings

99.98%
gemma-4-31B-it at N = 1 → 419 needles (two homograph keys excluded)
6.0×
more misses when needles sit over 1,024 tokens apart (12B)
13
Gemma and Qwen models compared on the same grid
261K
tokens of context at the longest setting
  1. gemma-4-31B-it shows no needle-count limit. From N = 1 to N = 419 it binds 21,227 of 21,252 needles (99.88%), with no downward trend. Excluding the two needle keys that are ordinary Turkish words, 99.98%. The 419-city key pool ran out before the model did.
  2. Every Gemma-4 model beats every Gemma-3 model on city-bound retrieval. Gemma-3 models find the right numbers but attach them to the wrong cities — gemma-3-4b-it loses 106 needles that value-only scoring would have credited, 96 of them written against another requested city; no Gemma-4 model does that more than 2 times.
  3. Qwen matches or beats the Gemma-4 model of similar size. Qwen3.8-27B binds all 420 needles, tying gemma-4-31B-it as the only perfect models. The Qwen3.6-35B-A3B mixture of experts (3B active) binds 99.8% against gemma-4-26B-A4B-it's 90.0% (4B active); Qwen3.5-9B 97.9% against gemma-4-12B-it's 92.4%; Qwen3.5-4B 95.2% against gemma-4-E4B-it's 83.6%; Qwen3.5-2B 96.9% against gemma-4-E2B-it's 97.1%. Of Qwen3.5-4B's 20 misses, 9 are the right number next to a Turkish or misspelled city name (BEYRUT for Beirut).
  4. Needle spacing, not needle count, predicts failure. gemma-4-12B-it misses 1.61% of needles packed within 1,024 tokens of each other and 9.68% when they sit further apart (6.0×). Needle count alone has no significant effect (ρ = -0.109, p = 0.10). Raising N at a fixed context length packs needles closer and makes the task easier.
  5. Needle keys must be screened against the haystack's language. Kabul is the everyday Turkish word for “acceptance” and occurs 86 times in the corpus; it caused 20 of the 31B's 25 misses (28.6% miss rate vs 0.014% for every other city). Asking gemma-4-31B-it about all 419 keys flagged 6 as Turkish words: Amman, Bari, Havana, Kabul, Kazan, Nice.

Methodology

Serving: vLLM 0.29.0 on a single NVIDIA A100-SXM4-80GB, temperature 0, bfloat16 (gemma-4-31B-it at 256K uses fp8 weights — the only way it fits). Qwen models were served text-only (--language-model-only) with thinking disabled (chat_template_kwargs: {"enable_thinking": false}): the Gemma models answer without reasoning, and the answer budget has no room for it. Haystack: 218 Turkish Wikipedia articles (trwiki-67), 200,210 Gemma tokens, extended to 238 articles for the 256K runs with the first 218 byte-identical. Grid: 15 context lengths, evenly spaced from 1,000 to 127,000 tokens (or 261,000). Cells too small to hold their needles plus 500 tokens of haystack are skipped (·). Context length is counted in each model's own tokens: Qwen's tokenizer averages 3.24 characters per token on this corpus against Gemma's 3.42, so a Qwen prompt of the same length holds about 5% less text.

Scoring: every multi-needle answer is re-scored from the stored response. A needle is retrieved only if its number appears on a line naming its own city; the parser accepts any reasonable line format and folds Turkish diacritics.

Needle and question (Turkish, as sent)

{city} şehrinin özel sihirli sayısı: {rnd_number}
Aşağıdaki şehirlerin özel sihirli sayıları nedir: {cities}? Her şehir için bulduğun sayıyı listele.

“The special magic number of the city of {city}: {number}” — “What are the special magic numbers of the following cities: {cities}? List the number you found for each city.”

Full prompt

Kullanıcıların sorularını yanıtlayan yardımcı bir yapay zeka botusunuz.
Aşağıda bir bağlam ve bağlamla ilgili bir soru bulunmaktadır.
#BAĞLAM
{context}
#BAĞLAM SONU

#SORU
{question} Belge dışından bilgi vermeyin. Yanıtını tam olarak şu biçimde ver:
SEHIR: <şehir adı> = <sayı>
(her şehir için bir satır, bulamadığın şehir için <sayı> yerine YOK yaz)
Başka hiçbir şey yazma.

“You are a helpful AI bot that answers users' questions. Below is a context and a question about it. … Do not give information from outside the document. Answer exactly in this format: CITY: <city> = <number> (one line per city; write YOK instead of the number for a city you cannot find). Write nothing else.”

Benchmark version

All results shown use 3-digit values, a single fixed needle sentence, and city keys that were not screened against Turkish vocabulary (see Known limitations).

Known limitations

  • Needle keys were not screened. The key pool includes Kabul, the everyday Turkish word for “acceptance”; its misses are counted, not excluded, in every table.
  • Qwen models answered with thinking disabled. That matches how the Gemma models answer, so the comparison is like for like, but it is not Qwen's default mode; with reasoning on they may score differently. Their context lengths are in Qwen tokens, which cover about 5% less text than Gemma tokens.
  • One run per cell, no repeated seeds. Repeating one condition (gemma-4-12B-it, N = 40, span 100%) gave 6/14 vs 4/14 perfect cells, so treat single-cell differences as noise.
  • gemma-4-31B-it at 256K is fp8, not bf16.

Data

Every CSV holds one row per call, including the model's raw response, re-scored with the city-bound scorer used on this page. Summaries are the tables above.