Olympus core trades clock for branch-target reach
NVIDIA's Olympus is a 10-wide, SMT server core at 3.3 GHz, and its two taken branches per cycle hold only under 48 KB instruction footprints.

3.3 GHz is the clock speed listed for Olympus. That number does not prove the core can win single-threaded performance, because it says nothing about IPC, power, or the branch predictor's behavior under real workloads.
Die Brief reads the document as a constraint on NVIDIA's server core: a 10-wide, SMT core with deep reordering must hide mispredicts, and the branch-target path shows where it pays. Two taken branches per cycle hold only while instruction footprints stay under 48 KB; beyond that, the 16K-entry BTB at 4-cycle latency takes over. The partitioning of branch-target capacity between two threads also cuts per-thread reach, so SMT is not a free lunch for branch-heavy server code. For a buyer, that means Olympus is strongest in tight, branch-heavy code, not in sprawling instruction streams.
The document leaves out process node, SKU, wattage, exact SPEC CPU2026 scores, release date, and IPC at 3.3 GHz. Those gaps keep the core's efficiency and deployment context open.
Olympus is a wide, SMT core whose branch-target path favors tight instruction footprints over sprawling code. For buyers, the constraint is real: two threads split the BTB, and 48 KB is where the fast path ends.
The document does not state process node, SKU, wattage, exact SPEC CPU2026 scores, release date, or IPC at 3.3 GHz.
The 10-wide, SMT design depends on hiding mispredicts, but the 48 KB instruction-footprint limit and 16K-entry, 4-cycle BTB define the fallback. SMT halves per-thread branch-target capacity, so branch-heavy workloads lose reach.
3.3 GHz clock speed · 10-wide out-of-order core · two taken branches per cycle · 48 KB instruction footprint threshold · 16K entry BTB · 4 cycle latency
After chipsandcheese.com. We did not report this. The pictures, if any, are theirs.


