Show HN: DeepSeek-V4 Latent Reasoning – moving "thinking" into latent space
DeepSeek-V4-Flash-0731-Latent-Reasoning is a self-contained model that aims to compress 'thinking' into latent representations rather than emitting thinking tokens in the output. Published on blog.n.ichol.ai and hosted on HuggingFace at nmitchko/DeepSeek-V4-Flash-0731-Latent-Reasoning, the model is built on the DeepSeek-V4-Flash-0731 backbone quantized down to NVFP4, with a production vllm form for serving runtime.
The model preserves the DSpark draft block (3 layers) from the source and adds a trained latent reasoning head of 35.7M parameters, loaded from a single latent_reasoning_head.safetensors file (~152 MB). The backbone weights are roughly 79 GiB per GPU at TP=2 (158–164 GiB total).
Initial evaluations were run on BIG-Bench Hard (cot_zeroshot, 27 subtasks, 50 items per subtask, 1350 items total) using lm-evaluation-harness 0.4.12 against an OpenAI-compatible endpoint. The aggregate flexible-extract score is 0.94 ± 0.008, while the exact-match aggregate is 0.880. Because the model does not emit the literal phrase 'The answer is X', strict-match regex scoring produces near-zero results that are formatting artifacts rather than reasoning failures; lm-eval uses flexible-extract. Per-subtask values carry about ±0.05–0.07 error at 50 items, so the aggregate of 0.880 is the reliable number.
The model is strongest on multi-step state tracking: tracking_shuffled_objects_three_objects, tracking_shuffled_objects_five_objects, tracking_shuffled_objects_seven_objects, boolean_expressions, formal_fallacies, and penguins_in_a_table all score 1.00. Other strong scores include word_sorting 0.98, temporal_sequences 0.98, object_counting 0.98, navigate 0.98, logical_deduction_three_objects 0.98, and hyperbaton 0.96. It is weakest on mechanical, syntax-heavy jobs: dyck_languages is the clear outlier at 0.26, with disambiguation_qa at 0.58, causal_judgement 0.66, and geometric_shapes 0.74 as genuine weaknesses, not measurement artifacts.
The latent reasoning head works by taking layer 35's 4096-d hidden state, applying LayerNorm, then passing it through linear layers with SiLU activations (4096→2048, 2048→2048, 2048→2048) to produce [mu, log_sigma]. A separate 'stop_head' takes the same normalized hidden state through 4096→1024→1 to produce an end-of-reasoning signal. The mu vector forms a 1024-d latent, followed by another LayerNorm.