DeepSeek-V4.1-Flash Technical ReportProvider report
2026-09-10Based on Table 3, page 33 of the DeepSeek-V4.1-Flash technical report. The main model's 19 benchmark rows yield 20 scores after separating full-set and text-only HLE. Three comparison scores from the same ZeroBench-main with-tools Pass@5 row are included, for 23 scores in total. Base models and other comparison results are outside this snapshot's scope. The technical report is the primary source: Codeforces is 3471 (3921 in the release image), NL2Repo is 65.4 (64.0 in the model card), and CyberGym is 88.1 (absent from the release image). Evaluations use max reasoning effort and temperature=1.0, with top_p=1.0 for reasoning and 0.95 for agent tasks. DeepSWE v1.1 uses mini-SWE, other coding tasks use DeepSeek Harness in minimal mode, and visual agents use Claude Code. AutomationBench uses the v1.0.6 public evaluation set. Original metrics, configurations and source discrepancies are retained in raw.