Agents and people, sharing the same machine

The multi-department build

Fifty people at 8K context, twelve at 32K, with agents still running.

Built for

$48,000 – $54,000

Parts alone

$44,520 – $50,520

The difference is a flat $3,000 for assembly, operating system and serving-engine configuration, burn-in testing and warranty handling. Flat, not a percentage — a bigger GPU should not cost you more in labour. Component prices verified 2026-08-22.

What this one is for

  • Several models resident at once rather than swapped in and out
  • Agent workflows in the background, not competing with people for the GPU
  • Long-context work — contracts, transcripts, codebases — as the normal case

Where it stops

This is the only build where a serious agent fleet is a real conversation, and even here it depends entirely on context length. Agent counts quoted without a context length are marketing. We size against measured usage, from a pilot.

Every part, and what it costs

Verified against live US retailer listings on 2026-08-22. Nothing here is estimated. A component we could not price would be missing rather than guessed at.

Part Specification Cost
GPU Two NVIDIA RTX PRO 6000 Blackwell, 96 GB each 192 GB total. The single largest line in any build we sell. $29,000 – $34,000
CPU AMD Ryzen Threadripper PRO 9975WX, 32 cores First-party, frequently on back order. The 16-core 9955WX is $1,549. $3,899
Motherboard ASUS Pro WS TRX50-SAGE WIFI Two full x16 Gen5 slots, both populated here. $851
Memory 256 GB DDR5 ECC RDIMM System memory should comfortably exceed total VRAM. $7,920
Storage 3.84 TB NVMe, enterprise, power-loss protected Micron 7450 PRO class. $2,000 – $3,000
Power supply 1500 W on a 240 V circuit Two 600 W cards will not run on 120 V. This is a 240 V build. $350
Cooling Noctua NH-U14S TR5-SP6 Airflow planning matters more than the cooler at this power. $140
Chassis Fractal Define 7 XL Clearance verified per build before order. $240
Network Intel X550-T2 dual 10 GbE 25 Gb SFP28 is cheaper at $78 if you have the switch for it. $120
Parts subtotal $44,520 – $50,520
Build, configuration, burn-in and warranty handling $3,000
Total $48,000 – $54,000

Component prices moved sharply through 2026 — memory roughly doubled in the first quarter alone, and GPUs more than doubled over the year. A quote is good for the week it is written, and we re-price honestly rather than carrying an old number forward.

How it is built and configured

Serving engine
vLLM or SGLang, pipeline parallel across both cards
Model format
AWQ or FP8, with FP8 KV cache doubling effective concurrency
Memory
ECC RDIMM, four channels
Power
Past what a 120 V circuit carries. Planned with an electrician as part of the build
Network
10 Gb, with 25 Gb available

Why the user counts have a context length on them

A model's weights are fixed, but every person talking to it needs their own context, and that is what fills the rest of the card. For a 70B model it costs about 320 KiB per token — so one person at 8K tokens uses 2.5 GiB and the same person reading a long contract at 32K uses 10 GiB. The same machine serves very different numbers of people depending on the work.

That is why every figure here names the context it assumes. A vendor quoting you a user count without one has not done the arithmetic.

See the full VRAM arithmetic

Why this card and not a cheaper one

Two cards with the same memory can answer at very different speeds, because for one person at a time generation tracks memory bandwidth more than anything else. Every card we costed, with the number most vendors leave off the table:

Card VRAM Bandwidth Power Price
RTX PRO 4000 Single slot and frugal, but 24 GB does not hold a 70B and 672 GB/s is slow to answer. 24 GB 672 GB/s 145 W $2,799–4,536
RTX PRO 4500 Same memory as a 5090 at half the bandwidth and a similar price. ECC and 200 W are real, but you feel the missing bandwidth on every reply. 32 GB 896 GB/s 200 W $3,999–5,499
RTX 5090 The bandwidth-per-dollar winner, by nearly double. What the small-team and small-business builds use. 32 GB 1,792 GB/s 575 W $4,300–5,000
RTX PRO 5000, 48 GB Enough for a 70B with little room left for people. An awkward middle. 48 GB 1,344 GB/s 300 W $6,500–8,699
RTX PRO 5000, 72 GB The cheapest memory on the table at about $119/GB. Worth asking us about if you want capacity over speed. 72 GB 1,344 GB/s 300 W $8,600–13,136
RTX PRO 6000, 96 GB In this build Same bandwidth as an RTX 5090 with three times the memory and ECC. You are buying room, not speed. 96 GB 1,792 GB/s 600 W $14,500–17,000

The RTX PRO 4500 is the one worth pointing at. Same 32 GB as an RTX 5090, error-correcting memory, a third of the power, and a blower designed to sit beside another card — at half the bandwidth for a similar price. On a spec sheet it looks like the smarter buy. In use it answers at roughly half the speed. We nearly specified it ourselves.

Tell us what you are trying to run

Describe the workload, how many people need it, and how sensitive the data is. You will get a straight answer about whether owning the hardware makes sense for that situation — including when it does not, and a subscription would serve you better.

Prefer to talk? Call James on 832-338-2926. Please do not send confidential, client-privileged, health, financial-account or credential information through this form.

Tell us what you are trying to run

Describe the workload, how many people need it, and how sensitive the data is. You will get a straight answer about whether owning the hardware makes sense for that situation — including when it does not, and a subscription would serve you better.

Prefer to talk? Call James on 832-338-2926. Please do not send confidential, client-privileged, health, financial-account or credential information through this form.

Call James Send details