Skip to content
coldsnap
Measured, not modeled

Startup results with the boundaries attached.

ColdSnap benchmarks report time to the first non-empty streamed token and require the exact validation response from every accepted sample.

Strongest accepted case6.9× fastervLLM · Qwen3.8 27B FP8
Accepted matrix

Accepted August 2026 qualification matrix

Median of three runs with Linux page cache cleared before every sample. TTFT runs from Docker container start to the first non-empty streamed token.

Strongest result6.9×vanilla time ÷ ColdSnap time
Validation36 / 36accepted restore samples
Hardware2 hostsNVIDIA DGX Spark GB10
EngineModelDriver / modeVanillaColdSnapSpeedup
vLLMQwen3.8 27B FP8n610 native229.914s33.149s6.9×
SGLangQwen3.8 27B FP8n610 native127.969s48.975s2.6×
vLLMDeepSeek V4 Flash 0731n610 native298.782s43.023s6.9×
These are selected n610 native headline cases, not universal expectations.
Results are qualified per engine, model, runtime image, CUDA version, driver, snapshot driver, hardware, weight path, and topology.
Methodology

What the timer includes.

The benchmark is intentionally specific so the number can be reproduced and interpreted correctly.

Start

Container activation

The timer begins at Docker container State.StartedAt. Manager preparation, verification, and capsule pulls performed before that point are excluded.

Finish

First streamed token

The timer ends on the first non-empty streamed model token, after the serving workload is ready to respond—not at process restore or a health check.

Discipline

Cold page cache

Linux page cache is cleared on both hosts before each sample. Reported values are three-sample medians and every accepted sample must validate.

Latest run

v0.3.18 qualification

Functional qualification complete; 10 of 12 ColdSnap restore performance rows passed.

The September 4 matrix expanded functional evidence while retaining two DeepSeek native performance follow-ups. Functional correctness and performance guardrails are reported separately.