A whole small company on one machine
The small business build
Eleven people at 4K context, five or six at 8K, on a 70B model.
Built for
$16,000 – $17,500
Parts alone
$12,899 – $14,299
The difference is a flat $3,000 for assembly, operating system and serving-engine configuration, burn-in testing and warranty handling. Flat, not a percentage — a bigger GPU should not cost you more in labour. Component prices verified 2026-08-22.
What this one is for
- A 70B-class model serving everyone in the company through a browser
- Document analysis and retrieval over your own files, alongside chat
- Replacing per-seat AI subscriptions with hardware you own outright
Where it stops
Two cards at 64 GB hold the model and roughly 14 GiB of context between them. That is eleven people on short questions and fewer than two on 32K-token documents. If long documents are the daily job rather than the exception, the medium-business build is the honest answer.
Every part, and what it costs
Verified against live US retailer listings on 2026-08-22. Nothing here is estimated. A component we could not price would be missing rather than guessed at.
| Part | Specification | Cost |
|---|---|---|
| GPU | Two NVIDIA RTX 5090, 32 GB each New. 64 GB total. Neither card has NVLink — consumer Blackwell dropped it. | $8,600 – $10,000 |
| CPU | AMD Ryzen 9 9950X, 16 cores AM5 drives both cards at PCIe 5.0 x8. Threadripper is not required here. | $549 |
| Motherboard | ASUS ProArt X870E-CREATOR WIFI Documented x8/x8 across both slots. The second M.2 socket stays empty, or slot two drops to x4. | $510 |
| Memory | 128 GB DDR5, 2 × 64 GB ECC UDIMM is $3,397 if you want it. We quote both. | $1,890 |
| Storage | 4 TB PCIe Gen4 NVMe WD Black SN850X, first-party listing. | $640 |
| Power supply | Corsair HX1500i, 1500 W, ATX 3.1 Chosen because it delivers its rating on 120 V. Most 2000 W units do not. | $350 |
| Cooling | Noctua NH-D15 class air Two cards make their own heat. Case airflow matters more than the cooler here. | $120 |
| Chassis | Fractal Define 7 XL Full tower, clearance for two triple-slot cards. | $240 |
| Parts subtotal | $12,899 – $14,299 | |
| Build, configuration, burn-in and warranty handling | $3,000 | |
| Total | $16,000 – $17,500 | |
Component prices moved sharply through 2026 — memory roughly doubled in the first quarter alone, and GPUs more than doubled over the year. A quote is good for the week it is written, and we re-price honestly rather than carrying an old number forward.
How it is built and configured
- Serving engine
- vLLM with pipeline parallelism
- Why not tensor parallel
- Neither card has NVLink, and vLLM documents pipeline parallelism as the faster choice without it
- Model format
- AWQ 4-bit, 37 GiB of weights
- Power
- About 1,500 W under load. Needs its own 20 A circuit
- Network
- 10 Gb onboard
Why the user counts have a context length on them
A model's weights are fixed, but every person talking to it needs their own context, and that is what fills the rest of the card. For a 70B model it costs about 320 KiB per token — so one person at 8K tokens uses 2.5 GiB and the same person reading a long contract at 32K uses 10 GiB. The same machine serves very different numbers of people depending on the work.
That is why every figure here names the context it assumes. A vendor quoting you a user count without one has not done the arithmetic.
See the full VRAM arithmeticWhy this card and not a cheaper one
Two cards with the same memory can answer at very different speeds, because for one person at a time generation tracks memory bandwidth more than anything else. Every card we costed, with the number most vendors leave off the table:
| Card | VRAM | Bandwidth | Power | Price |
|---|---|---|---|---|
| RTX PRO 4000 Single slot and frugal, but 24 GB does not hold a 70B and 672 GB/s is slow to answer. | 24 GB | 672 GB/s | 145 W | $2,799–4,536 |
| RTX PRO 4500 Same memory as a 5090 at half the bandwidth and a similar price. ECC and 200 W are real, but you feel the missing bandwidth on every reply. | 32 GB | 896 GB/s | 200 W | $3,999–5,499 |
| RTX 5090 In this build The bandwidth-per-dollar winner, by nearly double. What the small-team and small-business builds use. | 32 GB | 1,792 GB/s | 575 W | $4,300–5,000 |
| RTX PRO 5000, 48 GB Enough for a 70B with little room left for people. An awkward middle. | 48 GB | 1,344 GB/s | 300 W | $6,500–8,699 |
| RTX PRO 5000, 72 GB The cheapest memory on the table at about $119/GB. Worth asking us about if you want capacity over speed. | 72 GB | 1,344 GB/s | 300 W | $8,600–13,136 |
| RTX PRO 6000, 96 GB Same bandwidth as an RTX 5090 with three times the memory and ECC. You are buying room, not speed. | 96 GB | 1,792 GB/s | 600 W | $14,500–17,000 |
The RTX PRO 4500 is the one worth pointing at. Same 32 GB as an RTX 5090, error-correcting memory, a third of the power, and a blower designed to sit beside another card — at half the bandwidth for a similar price. On a spec sheet it looks like the smarter buy. In use it answers at roughly half the speed. We nearly specified it ourselves.
Tell us what you are trying to run
Describe the workload, how many people need it, and how sensitive the data is. You will get a straight answer about whether owning the hardware makes sense for that situation — including when it does not, and a subscription would serve you better.
Prefer to talk? Call James on 832-338-2926. Please do not send confidential, client-privileged, health, financial-account or credential information through this form.