DeepLearning.AI details DeepSeek cache reduction and Flash benchmark results in The Batch
DeepLearning.AI published a [technical breakdown](https://hubs.la/Q04ztt000) of DeepSeek in this week's issue of The Batch, focusing on memory footprint optimizations for AI agents. The report states that DeepSeek reduced its cache to 890 bytes per token, making it 437 times smaller than DeepSeek-V1.
The analysis also notes a 25% increase in compute per output token when scaling context from 4K to 1M input tokens. On benchmark evaluation, Flash scored 39 compared to 36 for V4-Pro on the AA Index, with costs reported at $0.27 per task versus $0.67 per task.