A whole small company on one machine

The small business build

Eleven people at 4K context, five or six at 8K, on a 70B model.

Built for

$16,000 – $17,500

Parts alone

$12,899 – $14,299

The difference is a flat $3,000 for assembly, operating system and serving-engine configuration, burn-in testing and warranty handling. Flat, not a percentage — a bigger GPU should not cost you more in labour. Component prices verified 2026-08-22.

What this one is for

  • A 70B-class model serving everyone in the company through a browser
  • Document analysis and retrieval over your own files, alongside chat
  • Replacing per-seat AI subscriptions with hardware you own outright

Where it stops

Two cards at 64 GB hold the model and roughly 14 GiB of context between them. That is eleven people on short questions and fewer than two on 32K-token documents. If long documents are the daily job rather than the exception, the medium-business build is the honest answer.

Every part, and what it costs

Verified against live US retailer listings on 2026-08-22. Nothing here is estimated. A component we could not price would be missing rather than guessed at.

Part Specification Cost
GPU Two NVIDIA RTX 5090, 32 GB each New. 64 GB total. Neither card has NVLink — consumer Blackwell dropped it. $8,600 – $10,000
CPU AMD Ryzen 9 9950X, 16 cores AM5 drives both cards at PCIe 5.0 x8. Threadripper is not required here. $549
Motherboard ASUS ProArt X870E-CREATOR WIFI Documented x8/x8 across both slots. The second M.2 socket stays empty, or slot two drops to x4. $510
Memory 128 GB DDR5, 2 × 64 GB ECC UDIMM is $3,397 if you want it. We quote both. $1,890
Storage 4 TB PCIe Gen4 NVMe WD Black SN850X, first-party listing. $640
Power supply Corsair HX1500i, 1500 W, ATX 3.1 Chosen because it delivers its rating on 120 V. Most 2000 W units do not. $350
Cooling Noctua NH-D15 class air Two cards make their own heat. Case airflow matters more than the cooler here. $120
Chassis Fractal Define 7 XL Full tower, clearance for two triple-slot cards. $240
Parts subtotal $12,899 – $14,299
Build, configuration, burn-in and warranty handling $3,000
Total $16,000 – $17,500

Component prices moved sharply through 2026 — memory roughly doubled in the first quarter alone, and GPUs more than doubled over the year. A quote is good for the week it is written, and we re-price honestly rather than carrying an old number forward.

How it is built and configured

Serving engine
vLLM with pipeline parallelism
Why not tensor parallel
Neither card has NVLink, and vLLM documents pipeline parallelism as the faster choice without it
Model format
AWQ 4-bit, 37 GiB of weights
Power
About 1,500 W under load. Needs its own 20 A circuit
Network
10 Gb onboard

Why the user counts have a context length on them

A model's weights are fixed, but every person talking to it needs their own context, and that is what fills the rest of the card. For a 70B model it costs about 320 KiB per token — so one person at 8K tokens uses 2.5 GiB and the same person reading a long contract at 32K uses 10 GiB. The same machine serves very different numbers of people depending on the work.

That is why every figure here names the context it assumes. A vendor quoting you a user count without one has not done the arithmetic.

See the full VRAM arithmetic

Why this card and not a cheaper one

Two cards with the same memory can answer at very different speeds, because for one person at a time generation tracks memory bandwidth more than anything else. Every card we costed, with the number most vendors leave off the table:

Card VRAM Bandwidth Power Price
RTX PRO 4000 Single slot and frugal, but 24 GB does not hold a 70B and 672 GB/s is slow to answer. 24 GB 672 GB/s 145 W $2,799–4,536
RTX PRO 4500 Same memory as a 5090 at half the bandwidth and a similar price. ECC and 200 W are real, but you feel the missing bandwidth on every reply. 32 GB 896 GB/s 200 W $3,999–5,499
RTX 5090 In this build The bandwidth-per-dollar winner, by nearly double. What the small-team and small-business builds use. 32 GB 1,792 GB/s 575 W $4,300–5,000
RTX PRO 5000, 48 GB Enough for a 70B with little room left for people. An awkward middle. 48 GB 1,344 GB/s 300 W $6,500–8,699
RTX PRO 5000, 72 GB The cheapest memory on the table at about $119/GB. Worth asking us about if you want capacity over speed. 72 GB 1,344 GB/s 300 W $8,600–13,136
RTX PRO 6000, 96 GB Same bandwidth as an RTX 5090 with three times the memory and ECC. You are buying room, not speed. 96 GB 1,792 GB/s 600 W $14,500–17,000

The RTX PRO 4500 is the one worth pointing at. Same 32 GB as an RTX 5090, error-correcting memory, a third of the power, and a blower designed to sit beside another card — at half the bandwidth for a similar price. On a spec sheet it looks like the smarter buy. In use it answers at roughly half the speed. We nearly specified it ourselves.

Tell us what you are trying to run

Describe the workload, how many people need it, and how sensitive the data is. You will get a straight answer about whether owning the hardware makes sense for that situation — including when it does not, and a subscription would serve you better.

Prefer to talk? Call James on 832-338-2926. Please do not send confidential, client-privileged, health, financial-account or credential information through this form.

Tell us what you are trying to run

Describe the workload, how many people need it, and how sensitive the data is. You will get a straight answer about whether owning the hardware makes sense for that situation — including when it does not, and a subscription would serve you better.

Prefer to talk? Call James on 832-338-2926. Please do not send confidential, client-privileged, health, financial-account or credential information through this form.

Call James Send details