Small Startup (5–15 people, budget-conscious, low query volume)
You need an in-house AI coding/writing assistant your team can use privately, without sending code to a third party — but you don't have a data-center budget.
Our Recommendation
Skip GPUs entirely for the big model. Instead, put a huge amount of ordinary (cheap) system RAM in one server and run a 4-bit compressed version of Kimi K2.7 Code on the CPU. It will answer more slowly than a GPU would, but for a small team's occasional use, that's a fair trade for a much lower bill.
Model: Kimi K2.7 Code (4-bit compressed)
Memory math: 1,000B params × 0.5 GB (4-bit is about half the size of the 1GB/B baseline) × 1.2 buffer = 600 GB of memory needed — fits comfortably inside 1 TB of RAM.
Memory math: 1,000B params × 0.5 GB (4-bit is about half the size of the 1GB/B baseline) × 1.2 buffer = 600 GB of memory needed — fits comfortably inside 1 TB of RAM.
Recommended Build
| Qty | Item | Unit Price | Subtotal |
|---|---|---|---|
| 1 | Dual EPYC server, 1 TB RAM configuration | $22,000 | $22,000 |
| 1 | RTX 4090 (handles the UI, embeddings, and small quick tasks) | $2,500 | $2,500 |
| Total | $24,500 | ||
Total power draw: 1,350 W
— see the Power & Watts page for what that means in real-world terms.
Request this build →