Mid-Size Company (200–1,000 people, customer-facing product, needs speed)
You're serving an AI feature to real paying customers around the clock and need fast, reliable answers at real-time chat speed — slow CPU inference isn't acceptable.
Our Recommendation
Build a real two-node GPU cluster of H100s connected with a proper InfiniBand switch, so the 16 GPUs act as one coordinated pool of memory and compute. This comfortably hosts GLM 5.2 Max at full quality with headroom for many simultaneous users.
Model: GLM 5.2 Max (full precision-equivalent)
Memory math: 753B params × 1 GB × 1.2 buffer = 903.6 GB needed. Two 8-GPU H100 nodes = 16 × 80 GB = 1,280 GB of combined GPU memory — enough to fit the model with about 376 GB left over for user traffic.
Memory math: 753B params × 1 GB × 1.2 buffer = 903.6 GB needed. Two 8-GPU H100 nodes = 16 × 80 GB = 1,280 GB of combined GPU memory — enough to fit the model with about 376 GB left over for user traffic.
Recommended Build
| Qty | Item | Unit Price | Subtotal |
|---|---|---|---|
| 2 | 8-GPU rack server chassis | $45,000 | $90,000 |
| 16 | NVIDIA H100 SXM5 80GB GPU | $28,000 | $448,000 |
| 1 | InfiniBand NDR switch (connects the two nodes into one real cluster) | $95,000 | $95,000 |
| Total | $633,000 | ||
Total power draw: 15,150 W
— see the Power & Watts page for what that means in real-world terms.
Request this build →