ByteDance Finds Position-Sensitive Weakness in DeepSeek V4
Seed researchers found DeepSeek V4 retrieval accuracy can vary sharply when the same fact moves within a long context, exposing a hidden cost of cache compression.
ByteDance’s Seed research team has identified a previously underappreciated weakness in DeepSeek’s long-context processing: the model’s ability to retrieve information can depend heavily on where that information appears in the prompt.
The researchers tested DeepSeek V4 Flash, V4 Pro and V4.1 Flash while keeping the question and target passage unchanged and moving the passage through a context of roughly 128,000 tokens. In the base V4 Flash model, retrieval accuracy varied by as much as 40.2 percentage points depending on the target’s position. Post-training reduced the gap to 19.1 points, while V4.1 Flash narrowed it further to 6.1 points.
Seed attributes the effect to DeepSeek’s block-based key-value cache compression. The technique reduces memory and bandwidth requirements by compressing consecutive tokens into fewer cache entries, making long-context inference cheaper. But the position of a fact inside a compressed block can affect how faithfully the model later reconstructs or retrieves it. The researchers describe the recurring pattern as “phase sensitivity.”
The finding matters beyond DeepSeek. Average scores on long-context benchmarks can conceal periodic blind spots, especially when a test places information at only a few fixed positions. For production systems, that creates a reliability problem: a document may be unchanged while a small formatting difference alters where critical content lands.
Seed’s results also show that post-training can mitigate, but not automatically eliminate, the weakness. Long-context evaluation will need position sweeps and repeated retrieval tests, rather than a single placement of each fact. The broader lesson is that memory-saving inference optimizations can introduce structured failure modes that aggregate benchmarks miss.
Uncle Cat take
The important result is the 40.2-point swing, not the “DeepSeek glitch” framing: cache efficiency has become part of model reliability, and buyers should test placement sensitivity before trusting long documents.