A few people, one model, one card
The small team build
Eight people at 4K context, four at 8K, on a 32B model.
Built for
$11,500 – $12,000
Parts alone
$8,199 – $8,899
The difference is a flat $3,000 for assembly, operating system and serving-engine configuration, burn-in testing and warranty handling. Flat, not a percentage — a bigger GPU should not cost you more in labour. Component prices verified 2026-08-22.
What this one is for
- Drafting, summarising and rewriting, in a browser, for a handful of people
- Question-answering over your own documents at modest volume
- Proving the idea works before committing to a bigger machine
Where it stops
This build does not run a 70B model. The weights alone are 39.8 GiB at Q4_K_M and this card holds 32 GB, so it runs 32B-class models instead. They are capable, and they are smaller. If a 70B is the requirement, start at the small-business build.
Every part, and what it costs
Verified against live US retailer listings on 2026-08-22. Nothing here is estimated. A component we could not price would be missing rather than guessed at.
| Part | Specification | Cost |
|---|---|---|
| GPU | NVIDIA RTX 5090, 32 GB GDDR7 New. 1,792 GB/s, 575 W. Launched at $1,999 in January 2025. | $4,300 – $5,000 |
| CPU | AMD Ryzen 9 9950X, 16 cores Retail boxed, first-party listing. | $549 |
| Motherboard | ASUS ProArt X870E-CREATOR WIFI 10 Gb and 2.5 Gb networking onboard. Frequently on back order. | $510 |
| Memory | 128 GB DDR5, 2 × 64 GB DDR5-5600. This line roughly doubled in the first quarter of 2026 alone. | $1,890 |
| Storage | 2 TB PCIe Gen4 NVMe WD Black SN850X or Samsung 990 Pro. | $390 |
| Power supply | 1000 W, 80 PLUS Platinum, ATX 3.1 Sized for one card with headroom, not for a second. | $200 |
| Cooling | Noctua NH-D15 class air Air, not liquid. Fewer things to fail unattended. | $120 |
| Chassis | Fractal Define 7 XL Room for the card and for airflow around it. | $240 |
| Parts subtotal | $8,199 – $8,899 | |
| Build, configuration, burn-in and warranty handling | $3,000 | |
| Total | $11,500 – $12,000 | |
Component prices moved sharply through 2026 — memory roughly doubled in the first quarter alone, and GPUs more than doubled over the year. A quote is good for the week it is written, and we re-price honestly rather than carrying an old number forward.
How it is built and configured
- Serving engine
- vLLM, with continuous batching configured
- Model format
- AWQ 4-bit, roughly 4.5 bits per weight
- Operating system
- Ubuntu LTS with the CUDA toolkit
- Power
- About 850 W under load, inside a standard 15 A circuit
- Network
- 2.5 Gb and 10 Gb onboard
Why the user counts have a context length on them
A model's weights are fixed, but every person talking to it needs their own context, and that is what fills the rest of the card. For a 70B model it costs about 320 KiB per token — so one person at 8K tokens uses 2.5 GiB and the same person reading a long contract at 32K uses 10 GiB. The same machine serves very different numbers of people depending on the work.
That is why every figure here names the context it assumes. A vendor quoting you a user count without one has not done the arithmetic.
See the full VRAM arithmeticWhy this card and not a cheaper one
Two cards with the same memory can answer at very different speeds, because for one person at a time generation tracks memory bandwidth more than anything else. Every card we costed, with the number most vendors leave off the table:
| Card | VRAM | Bandwidth | Power | Price |
|---|---|---|---|---|
| RTX PRO 4000 Single slot and frugal, but 24 GB does not hold a 70B and 672 GB/s is slow to answer. | 24 GB | 672 GB/s | 145 W | $2,799–4,536 |
| RTX PRO 4500 Same memory as a 5090 at half the bandwidth and a similar price. ECC and 200 W are real, but you feel the missing bandwidth on every reply. | 32 GB | 896 GB/s | 200 W | $3,999–5,499 |
| RTX 5090 In this build The bandwidth-per-dollar winner, by nearly double. What the small-team and small-business builds use. | 32 GB | 1,792 GB/s | 575 W | $4,300–5,000 |
| RTX PRO 5000, 48 GB Enough for a 70B with little room left for people. An awkward middle. | 48 GB | 1,344 GB/s | 300 W | $6,500–8,699 |
| RTX PRO 5000, 72 GB The cheapest memory on the table at about $119/GB. Worth asking us about if you want capacity over speed. | 72 GB | 1,344 GB/s | 300 W | $8,600–13,136 |
| RTX PRO 6000, 96 GB Same bandwidth as an RTX 5090 with three times the memory and ECC. You are buying room, not speed. | 96 GB | 1,792 GB/s | 600 W | $14,500–17,000 |
The RTX PRO 4500 is the one worth pointing at. Same 32 GB as an RTX 5090, error-correcting memory, a third of the power, and a blower designed to sit beside another card — at half the bandwidth for a similar price. On a spec sheet it looks like the smarter buy. In use it answers at roughly half the speed. We nearly specified it ourselves.
Tell us what you are trying to run
Describe the workload, how many people need it, and how sensitive the data is. You will get a straight answer about whether owning the hardware makes sense for that situation — including when it does not, and a subscription would serve you better.
Prefer to talk? Call James on 832-338-2926. Please do not send confidential, client-privileged, health, financial-account or credential information through this form.