Gorgon Halo's 192 GB pool carries a 320B GLM at 20 tokens per second
Loaded at 4-bit UD_IQ4_XS into roughly 160 GB of a 192 GB unified LPDDR5X pool, the 320B-parameter GLM 5.3 Flash runs on AMD's Ryzen AI Max+ PRO 495 Gorgon Halo at 20 tokens per second.

| Ryzen AI Max+ PRO 495 Gorgon Halo | Core Ultra X9 388H | |
|---|---|---|
| memory | 192 GB unified LPDDR5X | 64 GB |
| GPU memory assigned | 128 GB | |
| max memory | 192 GB | 96 GB |
| GLM 5.3 Flash | 320B parameters, 20 tokens/sec | |
| Qwen 3.8 Flash Next | 125B + 51B parameters, 42 tokens/sec | |
| ComfyUI leads | 18 of 18 | |
| lead range | 1.1x to 3.1x | |
| outlier | 32.2x |
The TechPowerUp write-up says the result came from one short physics prompt, so the 20 tokens per second figure does not establish sustained agentic throughput.
Memory assignment frames the comparison: AMD gave the GPU 128 GB for the ComfyUI comparison, while the Intel system had 64 GB.
The capacity gap is real, though. That makes the 18-workload comparison a capacity test as much as a performance test. The 32.2x Yuve lead is the clearest symptom, and the smaller 1.1x to 3.1x leads still leave the buyer without a per-workload cost figure.
No page measured token latency, power draw, or long-horizon agent task completion, so the buyer still lacks the cost-per-task number that would justify the 192 GB configuration.
The 20 tokens per second figure is a single-prompt result, and the 192 GB pool makes the comparison a capacity test rather than a clean silicon race.
After bsky.app. We did not report this. The pictures, if any, are theirs.

