IQuest-Q1's 320B MoE leaves self-hosting unpriced
IQuest-Q1 is a public 320B-parameter sparse MoE model with about 15B active per token, but the release leaves context length, latency, and hardware vendor unmeasured.
| IQuest-Q1 | |
|---|---|
| total parameters | ~320B |
| active per token | ~15B |
| architecture | Decoder-only Transformer with sparse Mixture-of-Experts feed-forward |
| training stages | three |
| benchmarks | NL2Repo, CyberGym, Terminal-Bench 2.1, DeepSWE v1.1, JobBench |
IQuest-Q1 is a decoder-only Transformer with a sparse Mixture-of-Experts feed-forward, about 320B total parameters and about 15B active per token. The weights and technical materials are public on GitHub, Hugging Face, and a blog. It is trained for repository navigation, terminal driving, tool calls, and multi-step tasks. The limit is self-hosting a model. The measurement is ~15B active per token. That active count sits below the total parameter count, but the release does not show it inside any named hardware budget.
Die Brief reads the constraint as buyer-side: the MoE design reduces per-token compute, but the 320B footprint still forces a self-hosting decision before benchmark claims matter. The release names integration with Claude Code and Codex CLI, but it does not state a supported stack. For a fab or cloud buyer, the active parameter count is the useful number, not the total. The document proves architecture and public weights, not latency, context length, or a deployment path.
What remains unmeasured is the runtime envelope: context length, inference latency, and hardware vendor are absent. The benchmark names are listed, but full numbers, baselines, and evaluation setup stay in the technical report. The one-shot app demos and RL debugging story are demonstrations, not measured throughput. No clock, power, price, or context window appears in the filing.
Die Brief reads the constraint as buyer-side: the MoE design reduces per-token compute, but the 320B footprint still forces a self-hosting decision before benchmark claims matter. The document proves architecture and public weights, not latency, context length, or a deployment path.
The release leaves context length, inference latency, and hardware vendor unmeasured.
The active 15B per token is the useful compute figure, while the 320B total sets the memory and self-hosting burden. Without context length, latency, or hardware vendor, the benchmark list is not a deployment spec.
After prnewswire.com. We did not report this. The pictures, if any, are theirs.

