Build tiers
Sized against a real workload, not a round number.
The reference build is Llama 3.3 70B, because it is a real model with published file
sizes you can check rather than a round number nobody can verify. Everything below is sized
against it. Starting prices for a machine you own outright, built from new parts
with the manufacturer's warranty intact.
These numbers are higher than they were a year ago, and we are not going to
pretend otherwise. GPU and memory prices have gone up
sharply through 2026 — the detail is below the tiers, with the
receipts.
Why 64 GB, and not less
Sized against Llama 3.3 70B — a real model with
published file sizes you can check, rather than a round number. Its weights alone,
before a single person connects:
- FP16
- 131.4 GiB
- Q8_0
- 69.8 GiB
- Q6_K
- 54.2 GiB
- Q4_K_M (what we ship)
- 39.8 GiB
Note that “4-bit” is never 4.0 bits. Q4_K_M
measures about 4.85, because the quantisation keeps attention and output layers at higher
precision. Plenty of vendors round it down to 35 GB. It is 39.8 GiB, and that gap is the
difference between a machine that fits your model and one that does not.
Then every person needs their own
Whatever is left after the weights goes to context, and context is per user, not shared.
For this model it costs 320 KiB per token:
- 4K context
- 1.25 GiB
- 8K context
- 2.5 GiB
- 32K context
- 10 GiB
- 128K context
- 40 GiB
This is why every concurrency number below names the context length it assumes. A machine
that serves eleven people on short questions serves fewer than two on long documents, on
exactly the same hardware. Anyone quoting you a user count without saying at what context
has not done the arithmetic.
one-time · owned outright
- VRAM
- 32 GB on one card
- GPU
- One RTX 5090 32 GB, new, around $4,300–5,000
- CPU
- Ryzen 9, 16 cores
- At once
- 8 at 4K context, 4 at 8K, on a 32B model
- Power
- About 850 W under load — fits a standard circuit
-
128 GB DDR5, 2 TB NVMe, 2.5 Gb networking
-
New hardware with the full manufacturer warranty
-
Linux, CUDA, vLLM, Open WebUI, configured and handed over
The catch
This tier does not run the 70B model the rest of this page is sized against. Its weights are 39.8 GiB at Q4_K_M and this card holds 32 GB, so it runs 32B-class models instead. Those are capable, and they are smaller. If a 70B is what you actually need, start at Business — we would rather lose the smaller sale than sell you a machine you outgrow in a month.
Electricity: about $160/year at Texas commercial rates.
See every part and price
Get My Custom Quote
Best value
one-time · owned outright
- VRAM
- 64 GB (2 × 32 GB)
- GPU
- Two RTX 5090 32 GB, around $4,300–5,000 each
- CPU
- Threadripper 9960X or 9970X, 24–32 cores
- At once
- 11 at 4K context, 5–6 at 8K, on a 70B
- Power
- About 1,500 W under load — needs a 20 A circuit or a 240 V drop
-
128–256 GB DDR5 ECC, 4 TB NVMe, 10 Gb networking
-
Everyone works in a browser — no workstation each
-
Linux, CUDA, vLLM or SGLang, Open WebUI, configured and handed over
The catch
Two 5090s draw 1,150 W between them, and a Threadripper adds roughly 350 W. Electrical code caps a continuous load at 80% of the breaker, which is 1,440 W on a standard 120 V 15 A circuit — so this build needs its own 20 A circuit or a 240 V drop. We check that before quoting, not after delivery.
Electricity: about $250/year at Texas commercial rates.
See every part and price
Get My Custom Quote
one-time · owned outright
- VRAM
- 96 GB on a single GPU
- GPU
- One RTX PRO 6000 Blackwell, 96 GB GDDR7 ECC
- CPU
- Threadripper, 32 cores
- At once
- 35 at 4K context, 17–18 at 8K, on a 70B
- Power
- About 1,000 W under load with the Max-Q card
-
Whole model and every user cache on one card
-
ECC memory, professional drivers, longer support life
-
No multi-GPU splitting to tune or troubleshoot
The catch
The GPU alone is $14,500–17,000 at August 2026 prices, up from $8,565 when it launched in March 2025 — about 87%, with no public explanation. It moved 21% in the first three weeks of this month alone. That single component is most of this tier, and it is the number most likely to have changed by the time you read this.
Electricity: about $190/year at Texas commercial rates.
See every part and price
Get My Custom Quote
one-time · owned outright
- VRAM
- 192 GB+
- GPU
- Dual professional GPUs, 96 GB each
- CPU
- Threadripper PRO, 32–64 cores
- At once
- 50 at 8K context, 12 at 32K, plus agents in the background
- Power
- Rack-class. Circuit and cooling planned as part of the build
-
Several models resident at once, not swapped in and out
-
Headroom for document analysis and RAG alongside chat
-
The tier where agent fleets stop competing with people for the GPU
The catch
This is the only tier where "120 agents" starts to be a real conversation, and even then it depends entirely on context length. At 8K context a measured dual-96GB machine sustained 125 concurrent; at 32K the same machine managed eight. We size against your actual context, not the headline.
Electricity: about $400/year at Texas commercial rates.
See every part and price
Get My Custom Quote
Why these prices have gone up
The parts these machines are made of have repriced hard, driven by demand that
has nothing to do with small business. What that has done, part by part:
-
RTX 5090 32 GB
$1,999 at launch → $4,300–5,000 today
· more than doubled
-
RTX PRO 6000 96 GB
$8,565 at launch → $14,500–17,000 today
· up ~87% since its March 2025 launch
-
DDR5 ECC memory
early-2026 pricing → roughly double that
· doubled in a quarter
We cannot tell you whether this settles or keeps climbing, and anyone who says
they can is guessing. What we can do is date every quote, show you the component
prices it was built from, and re-price honestly if the market moves before you
decide. A machine you own still stops costing money once it is paid off —
that part has not changed.
The comparison that actually matters
Chat seats are a flat fee. Agents are a meter.
Ten people on a chat subscription costs roughly
$200–$250 a month.
Against that alone, a $17,000 machine takes
five to seven years to pay
back. We would rather put that number in front of you than have you work it
out afterwards.
Agents change the arithmetic completely. They read far more than they write
— roughly 25 tokens in for every 1 out
— and input is about 85% of
the bill. Published analysis puts a
25-person team running a couple of
agent sessions each per day at around
$6,000 a month
in metered API spend. That is the number these machines actually stand
against.
What we are not claiming
-
That your bill is $6,000.
It depends on how many agents run, how long their context gets and how
much of it caches. Your usage decides it, not us.
-
That a 70B open model matches a frontier model. It does not. What you
get is unmetered usage, data that stays put, and a cost that stops
moving.
-
That owning is free to run. Electricity, support and eventual parts
continue, at roughly
$14–35 a month in Texas power.
Run it against your own numbers
The big models, and what they need
-
DeepSeek-V3 / R1 (671B)
about 407 GB at 4-bit — beyond every machine we build, flagship included
-
Llama 4 Maverick (400B)
about 245 GB at 4-bit — the flagship, and only just — almost nothing left for context
-
Qwen3-235B-A22B
about 142 GB at 4-bit — the flagship, comfortably
Anyone telling you a workstation runs full DeepSeek-R1 is rounding. It does not fit on
anything we build, flagship included, and we would rather you heard that here. The other
two need the flagship,
which is a different class of machine rather than a bigger version of these.
Concurrency figures assume vLLM or
SGLang. A default
Ollama install serves one request at a time, so the
same hardware behaves very differently depending on how it is configured. We ship it
configured.
Component prices verified 2026-08-21, in a market
that has moved sharply enough that a quote is only good for the week it is
written. Electricity estimated at
8.56¢/kWh, the Texas commercial
average (EIA, year to date through May 2026).