---
license: other
license_name: qwen-community-1.0
license_link: LICENSE
language: [en, ko, zh, ja, multilingual]
library_name: transformers
pipeline_tag: image-text-to-text
tags:
  - darwin
  - darwin-rsi
  - recursive-self-improvement
  - self-improvement
  - vidraft
  - final-bench
  - qwen
  - qwen3.8
  - moe
  - mixture-of-experts
  - sparse-moe
  - 180b
  - hybrid-attention
  - linear-attention
  - long-context
  - 262k-context
  - vision-language
  - multimodal
  - reasoning
  - reasoning-model
  - thinking
  - chain-of-thought
  - math
  - science
  - stem
  - ztc
  - zero-token-confidence
  - confidence-estimation
  - hallucination-detection
  - gpqa
  - gpqa-diamond
  - mmlu-pro
  - mmmu-pro
  - eval-results
  - korean
  - english
  - vllm
  - openai-compatible
  - b200
model-index:
  - name: Darwin-180B-RSI
    results:
      - task: {type: text-generation, name: Graduate-Level Reasoning}
        dataset: {type: Idavidrein/gpqa, name: GPQA Diamond, config: gpqa_diamond, split: train}
        metrics:
          - {type: accuracy, value: 94.44, name: "Accuracy (majority vote, up to 16 samples, 131K thinking)", verified: false}
      - task: {type: text-generation, name: Multi-discipline Knowledge & Reasoning}
        dataset: {type: TIGER-Lab/MMLU-Pro, name: MMLU-Pro, split: test}
        metrics:
          - {type: accuracy, value: 88.12, name: "Accuracy (single sample, 131K thinking)", verified: false}
      - task: {type: image-text-to-text, name: Multimodal Expert Reasoning}
        dataset: {type: MMMU/MMMU_Pro, name: MMMU-Pro (vision), config: vision, split: test}
        metrics:
          - {type: accuracy, value: 79.48, name: "Accuracy (majority vote, 3 samples, 131K thinking)", verified: false}
---

# Darwin-180B-RSI

### 180B Mixture-of-Experts · vision-language · **#1 on three Hugging Face official leaderboards** — GPQA Diamond 94.44 · MMLU-Pro 88.12 · MMMU-Pro 79.48 · **self-improving**

`reasoning` · `MoE 512 experts` · `262K long context` · `image + text` · `Korean + English` · `self-improvement` · `ZTC`

<p align="center">
<a href="https://vidraft.net"><img src="https://img.shields.io/badge/🌐_VIDRAFT-vidraft.net-111827?style=for-the-badge"></a>
<a href="https://huggingface.co/datasets/Idavidrein/gpqa"><img src="https://img.shields.io/badge/GPQA_Diamond-94.44%25_%231-gold?style=for-the-badge"></a>
<a href="https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro"><img src="https://img.shields.io/badge/MMLU--Pro-88.12%25_%231-2563eb?style=for-the-badge"></a>
<a href="https://huggingface.co/datasets/MMMU/MMMU_Pro"><img src="https://img.shields.io/badge/MMMU--Pro-79.48%25_%231-0891b2?style=for-the-badge"></a>
<img src="https://img.shields.io/badge/Self--Improving-RSI-e11d48?style=for-the-badge">
<img src="https://img.shields.io/badge/ZTC-Zero--Token_Confidence-7c3aed?style=for-the-badge">
</p>
<p align="center">
<a href="https://arxiv.org/abs/2605.14386"><img src="https://img.shields.io/badge/arXiv-2605.14386_Darwin_Family-b31b1b?style=for-the-badge"></a>
<a href="https://huggingface.co/papers/2609.20269"><img src="https://img.shields.io/badge/Paper-2609.20269_Latin_Square-b31b1b?style=for-the-badge"></a>
<a href="https://huggingface.co/collections/FINAL-Bench/darwin-family"><img src="https://img.shields.io/badge/🧬_Collection-Darwin_Family-16a34a?style=for-the-badge"></a>
<a href="https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems"><img src="https://img.shields.io/badge/🏛️_Collection-ZTC_Models-7c3aed?style=for-the-badge"></a>
</p>

**The newest flagship of the Darwin family — #1 on GPQA Diamond, MMLU-Pro and MMMU-Pro,
and a model that gets better by learning from its own verified work.**

---

## 🧬 The Darwin Family

<p align="center">
<a href="https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC"><img src="https://img.shields.io/badge/Darwin--397B--ZTC-GPQA_93.43-16a34a"></a>
<a href="https://huggingface.co/FINAL-Bench/Darwin-28B-REASON"><img src="https://img.shields.io/badge/Darwin--28B--REASON-GPQA_89.39-16a34a"></a>
<a href="https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus"><img src="https://img.shields.io/badge/Darwin--35B--A3B--Opus-♥98-e11d48"></a>
<a href="https://huggingface.co/FINAL-Bench/Darwin-36B-Opus"><img src="https://img.shields.io/badge/Darwin--36B--Opus-♥97-e11d48"></a>
<a href="https://huggingface.co/FINAL-Bench/Darwin-4B-Genesis"><img src="https://img.shields.io/badge/Darwin--4B--Genesis-♥63-e11d48"></a>
</p>
<p align="center">
<a href="https://huggingface.co/FINAL-Bench/Darwin-9B-NEG"><img src="https://img.shields.io/badge/Darwin--9B--NEG-♥57-e11d48"></a>
<a href="https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF"><img src="https://img.shields.io/badge/POCKET--35B-824K_↓-1f6feb"></a>
<a href="https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF"><img src="https://img.shields.io/badge/POCKET--26B-365K_↓-1f6feb"></a>
<a href="https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF"><img src="https://img.shields.io/badge/POCKET--EN-♥43-1f6feb"></a>
<a href="https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF"><img src="https://img.shields.io/badge/POCKET--KR-♥36-1f6feb"></a>
</p>

**Darwin** is [VIDRAFT](https://vidraft.net)'s measurement-driven reasoning model family —
**50+ official models**, **400+ community derivatives**, and now **two places in the GPQA Diamond top 3**
(Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).

---

## 🧬 Darwin — evolve the parent, keep what works

Darwin treats a strong open model as a **parent**. It measures where the parent is weak,
and strengthens exactly those parts — instead of re-training everything and risking what already works.

- **Diagnose before you change.** Every Darwin generation starts from a measured weakness map of the parent.
- **Change little, precisely.** Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
- **Proven capability over new guesses.** Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; Darwin-180B-RSI adds a new ingredient — **the model's own verified work**.
- **Measured, not claimed.** Every change must beat the parent on held-out tests before it ships.

| Model | Scale | GPQA Diamond |
|:---|:---|:---:|
| Darwin-9B-NEG | 9B | 84.3 |
| Darwin-27B-Opus | 27B dense | 86.9 |
| Darwin-36B-Opus | 36B MoE | 88.4 |
| Darwin-28B-REASON | 28B + DELPHI | 89.39 |
| Darwin-397B-ZTC | 397B MoE (FP8) | 93.43 |
| **Darwin-180B-RSI** | **180B MoE** | **94.44** |

### Lineage

| Role | | |
|:---|:---|:---|
| **Parent** | `Qwen/Qwen3.8-Flash-Next` | 180B MoE vision-language backbone · Qwen Community License 1.0 |
| **Darwin RSI** | self-improvement on verified answers | the parent's own solutions, checked against verifiable answer keys, fed back as training signal |
| **Preserved** | 512 routed experts · router · vision encoder | untouched — the parent's knowledge stays intact |
| **ZTC** | zero-token confidence readout | see below |

---

## 📄 Darwin Platform & Research

- **Darwin Family** — MRI trust-weighted evolutionary merging for training-free scaling of language-model reasoning ([arXiv:2605.14386](https://arxiv.org/abs/2605.14386))
- **Placement Is Free, Composition Is Not** — the Latin square as a provably-balanced construction for heterogeneous sequence-mixer stacks ([2609.20269](https://huggingface.co/papers/2609.20269)) — the AETHER architecture line
- **FINAL Bench** — VIDRAFT's measurement-driven evaluation framework (SSRN)
- **Four-layer Pre-AGI roadmap** — Darwin → AETHER → PROMETHEUS → HEPHAESTUS
- Collections: [Darwin Family](https://huggingface.co/collections/FINAL-Bench/darwin-family) · [ZTC Models — JEV ecosystems](https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems)

---

## 🔁 RSI — a model that improves from its own work

**Recursive self-improvement (RSI)** is the core of this generation.
Instead of distilling a bigger teacher, the model improves by learning from itself:

1. **Solve** — the model works through practice problems it has never seen in evaluation.
2. **Verify** — its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
3. **Learn** — it is re-trained on the reasoning that turned out to be correct.
4. **Repeat** — the improved model becomes the next solver.

What it bought in this release:

| | Parent (Qwen3.8-Flash-Next) | **Darwin-180B-RSI** |
|:---|:---:|:---:|
| Average reasoning length (MMLU-Pro) | 4,320 tokens | **3,833 tokens (−11 %)** |
| MMLU-Pro accuracy | 88.04 % | **88.12 %** |

**Same or better accuracy with shorter reasoning** — cheaper and faster to serve.
Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).

---

## 🏛️ ZTC — it knows before it answers

**Zero-Token Confidence (ZTC)** reads the model's own internal state **once, before generation**,
and returns the probability that the answer it is about to give is correct — **no extra tokens, no second model.**

```json
{"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false}
```

Use it to gate actions: when confidence is low, do not call the tool, escalate, or answer "I don't know".
The ZTC readout for this model is being fitted and will ship in `ztc/` (same format as
[Darwin-397B-ZTC](https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC)).

---

## 🏆 Results

| Benchmark | Score | Setting | Leaderboard |
|:---|:---:|:---|:---|
| **GPQA Diamond** (198) | **94.44** | majority vote over up to 16 samples · 131,072-token thinking budget | [**#1**](https://huggingface.co/datasets/Idavidrein/gpqa) |
| **MMLU-Pro** (12,032) | **88.12** | single sample · 131,072-token thinking budget | [**#1**](https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro) |
| **MMMU-Pro** (vision, 1,730) | **79.48** | majority vote over 3 samples · 131,072-token thinking budget | [**#1**](https://huggingface.co/datasets/MMMU/MMMU_Pro) |

Sampling for all runs: temperature 1.0 · top_p 0.95 · top_k 20 · bf16. All numbers are self-measured and reproducible with the settings above.

**MMLU-Pro by category (single sample)** — strongest in math 95.0 · biology 94.6 · physics 92.5; room to grow in law and history.

---

## ⚙️ Specifications

| | |
|:---|:---|
| Architecture | Mixture-of-Experts, hybrid attention (36 linear-attention + 12 full-attention layers) |
| Layers / hidden | 48 / 2,560 |
| Experts | 512 routed (10 active per token) + shared expert |
| Context | 262,144 tokens |
| Vocabulary | 248,320 |
| Modalities | image + text → text |
| Precision | bf16 (~336 GB) |

---

## 🚀 Quickstart

### Serving with vLLM (8 × B200 or equivalent)

```bash
vllm serve FINAL-Bench/Darwin-180B-RSI \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --max-model-len 135168 --trust-remote-code
```

### Chat Completions (OpenAI-compatible)

```python
from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI",
    messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
    temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
print(r.choices[0].message.content)
```

### Transformers

```python
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "FINAL-Bench/Darwin-180B-RSI"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
```

**Tip:** this is a thinking model. Give it room — a thinking budget of 32K–131K tokens is recommended for hard reasoning.
Short budgets truncate the reasoning and cost accuracy.

---

## ⚠️ Limitations and disclosure

- Scores are self-measured with the settings stated in the Results table; majority-vote numbers use several samples per question.
- Very long reasoning is normal for hard problems; a short thinking budget will truncate answers and lower accuracy.
- Like every LLM, the model can be confidently wrong — use the ZTC confidence readout to gate high-stakes actions.

---

## 🔗 Related Darwin Models

- **[Darwin-397B-ZTC](https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC)** — 397B MoE (FP8), GPQA Diamond 93.43 %, ZTC on board
- **[Darwin-28B-REASON](https://huggingface.co/FINAL-Bench/Darwin-28B-REASON)** — 28B, GPQA Diamond 89.39 %
- **[Darwin-36B-Opus](https://huggingface.co/FINAL-Bench/Darwin-36B-Opus)** — 36B MoE, GPQA Diamond 88.4 %
- **[Darwin-27B-RSI](https://huggingface.co/FINAL-Bench/Darwin-27B-RSI)** — 27B, the first Darwin RSI model
- **[Darwin-9B-NEG](https://huggingface.co/FINAL-Bench/Darwin-9B-NEG)** — 9B with Negentropy distillation, GPQA Diamond 84.3 %
- **[ZTC-Judge-27B](https://huggingface.co/FINAL-Bench/ZTC-Judge-27B)** — standalone ZTC judge

---

## 📚 Citation

```bibtex
@misc{darwin180b_rsi_2026,
  title  = {Darwin-180B-RSI: Recursive Self-Improvement on Verified Answers for a 180B Mixture-of-Experts Reasoning Model},
  author = {FINAL-Bench / Darwin Research Team},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-180B-RSI}},
  note   = {GPQA Diamond 94.44 \% · MMLU-Pro 88.12 \%}
}

@misc{darwin_family_2026,
  title  = {Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning},
  author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Shin, Junghoon},
  year   = {2026},
  eprint = {2605.14386},
  archivePrefix = {arXiv}
}

@misc{latin_square_2026,
  title  = {Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks},
  author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Kim, Minseo},
  year   = {2026},
  eprint = {2609.20269},
  archivePrefix = {arXiv}
}
```

---

## 📜 License

Darwin-180B-RSI is a derivative of **Qwen3.8-Flash-Next** and is distributed under the **Qwen Community License 1.0** (see `LICENSE`).

## 🏢 About

Built by **[VIDRAFT](https://vidraft.net)** · evaluated with **FINAL-Bench**.

This model is part of the [Darwin Family](https://arxiv.org/abs/2605.14386).
