Agents and people, sharing the same machine
The multi-department build
Fifty people at 8K context, twelve at 32K, with agents still running.
Built for
$48,000 – $54,000
Parts alone
$44,520 – $50,520
The difference is a flat $3,000 for assembly, operating system and serving-engine configuration, burn-in testing and warranty handling. Flat, not a percentage — a bigger GPU should not cost you more in labour. Component prices verified 2026-08-22.
What this one is for
- Several models resident at once rather than swapped in and out
- Agent workflows in the background, not competing with people for the GPU
- Long-context work — contracts, transcripts, codebases — as the normal case
Where it stops
This is the only build where a serious agent fleet is a real conversation, and even here it depends entirely on context length. Agent counts quoted without a context length are marketing. We size against measured usage, from a pilot.
Every part, and what it costs
Verified against live US retailer listings on 2026-08-22. Nothing here is estimated. A component we could not price would be missing rather than guessed at.
| Part | Specification | Cost |
|---|---|---|
| GPU | Two NVIDIA RTX PRO 6000 Blackwell, 96 GB each 192 GB total. The single largest line in any build we sell. | $29,000 – $34,000 |
| CPU | AMD Ryzen Threadripper PRO 9975WX, 32 cores First-party, frequently on back order. The 16-core 9955WX is $1,549. | $3,899 |
| Motherboard | ASUS Pro WS TRX50-SAGE WIFI Two full x16 Gen5 slots, both populated here. | $851 |
| Memory | 256 GB DDR5 ECC RDIMM System memory should comfortably exceed total VRAM. | $7,920 |
| Storage | 3.84 TB NVMe, enterprise, power-loss protected Micron 7450 PRO class. | $2,000 – $3,000 |
| Power supply | 1500 W on a 240 V circuit Two 600 W cards will not run on 120 V. This is a 240 V build. | $350 |
| Cooling | Noctua NH-U14S TR5-SP6 Airflow planning matters more than the cooler at this power. | $140 |
| Chassis | Fractal Define 7 XL Clearance verified per build before order. | $240 |
| Network | Intel X550-T2 dual 10 GbE 25 Gb SFP28 is cheaper at $78 if you have the switch for it. | $120 |
| Parts subtotal | $44,520 – $50,520 | |
| Build, configuration, burn-in and warranty handling | $3,000 | |
| Total | $48,000 – $54,000 | |
Component prices moved sharply through 2026 — memory roughly doubled in the first quarter alone, and GPUs more than doubled over the year. A quote is good for the week it is written, and we re-price honestly rather than carrying an old number forward.
How it is built and configured
- Serving engine
- vLLM or SGLang, pipeline parallel across both cards
- Model format
- AWQ or FP8, with FP8 KV cache doubling effective concurrency
- Memory
- ECC RDIMM, four channels
- Power
- Past what a 120 V circuit carries. Planned with an electrician as part of the build
- Network
- 10 Gb, with 25 Gb available
Why the user counts have a context length on them
A model's weights are fixed, but every person talking to it needs their own context, and that is what fills the rest of the card. For a 70B model it costs about 320 KiB per token — so one person at 8K tokens uses 2.5 GiB and the same person reading a long contract at 32K uses 10 GiB. The same machine serves very different numbers of people depending on the work.
That is why every figure here names the context it assumes. A vendor quoting you a user count without one has not done the arithmetic.
See the full VRAM arithmeticWhy this card and not a cheaper one
Two cards with the same memory can answer at very different speeds, because for one person at a time generation tracks memory bandwidth more than anything else. Every card we costed, with the number most vendors leave off the table:
| Card | VRAM | Bandwidth | Power | Price |
|---|---|---|---|---|
| RTX PRO 4000 Single slot and frugal, but 24 GB does not hold a 70B and 672 GB/s is slow to answer. | 24 GB | 672 GB/s | 145 W | $2,799–4,536 |
| RTX PRO 4500 Same memory as a 5090 at half the bandwidth and a similar price. ECC and 200 W are real, but you feel the missing bandwidth on every reply. | 32 GB | 896 GB/s | 200 W | $3,999–5,499 |
| RTX 5090 The bandwidth-per-dollar winner, by nearly double. What the small-team and small-business builds use. | 32 GB | 1,792 GB/s | 575 W | $4,300–5,000 |
| RTX PRO 5000, 48 GB Enough for a 70B with little room left for people. An awkward middle. | 48 GB | 1,344 GB/s | 300 W | $6,500–8,699 |
| RTX PRO 5000, 72 GB The cheapest memory on the table at about $119/GB. Worth asking us about if you want capacity over speed. | 72 GB | 1,344 GB/s | 300 W | $8,600–13,136 |
| RTX PRO 6000, 96 GB In this build Same bandwidth as an RTX 5090 with three times the memory and ECC. You are buying room, not speed. | 96 GB | 1,792 GB/s | 600 W | $14,500–17,000 |
The RTX PRO 4500 is the one worth pointing at. Same 32 GB as an RTX 5090, error-correcting memory, a third of the power, and a blower designed to sit beside another card — at half the bandwidth for a similar price. On a spec sheet it looks like the smarter buy. In use it answers at roughly half the speed. We nearly specified it ourselves.
Tell us what you are trying to run
Describe the workload, how many people need it, and how sensitive the data is. You will get a straight answer about whether owning the hardware makes sense for that situation — including when it does not, and a subscription would serve you better.
Prefer to talk? Call James on 832-338-2926. Please do not send confidential, client-privileged, health, financial-account or credential information through this form.