Sofiya's AI Hardware Store
Companies everywhere want to run big open-source AI models, and they all need NVIDIA hardware - but the spec sheets are confusing and the price tags are intimidating. This store sells real hardware, at real prices, and teaches you exactly what you're buying and why, before you spend a dollar.
Built and run by Sofiya Alterman.
The hardware ladder
Five tiers, from a desktop card to a full datacenter rack. Every tier shows its price, memory, and power draw - and what that means in plain English.
Desktop GPU
A single high-end graphics card you'd put in a normal desktop tower. This is the card gamers and hobbyists buy -- the same hardware, repurposed to run a smaller AI model on your desk.
- GPU Memory24 GB
- Power450 W
Representative photo
Professional Workstation GPU
A card built for professional workstations rather than gaming: more memory, ECC (error-correcting) memory for reliability, and certified drivers. It's the step up once a gaming card no longer has enough memory.
- GPU Memory96 GB
- Power600 W
Datacenter GPU
A card built for datacenters, not desks -- no video output, meant to sit in a rack and run flat-out on AI workloads. This is the hardware companies rent by the hour from cloud providers.
- GPU Memory80 GB
- Power350 W
Representative photo
Multi-GPU AI Server
A single purpose-built server with 8 datacenter GPUs already wired together with very fast direct connections (NVIDIA NVLink), so all 8 GPUs can share memory and work almost as if they were one giant GPU. You buy it as one complete machine, not 8 separate cards.
- GPU Memory640 GB
- Power7000 W
Representative photo
Full AI Datacenter Rack
An entire rack: 72 GPUs and 36 CPUs, wired together as one system with NVIDIA's fastest interconnects, delivered as a single unit. This is what a large AI company buys, not a single company usually -- and it's the top of the hardware ladder in this store.
- GPU Memory20700 GB
- Power135000 W
Three real models, sized honestly
We show the exact memory math for each model, so you can see for yourself why a given build is recommended - not just take our word for it.
Mixtral 8x22B
141 billion parameters (Mixture-of-Experts (8 experts, 2 active per token))
See the memory math →Qwen3.5-397B-A17B
397 billion parameters (Mixture-of-Experts (512 experts, 11 active per token))
See the memory math →DeepSeek-R1
671 billion parameters (Mixture-of-Experts (256 experts, 8 active per token))
See the memory math →What Does Your Model Need?
Type in any model's parameter count and get an instant estimate, using the same simplified formula shown throughout this store.
Educational estimate, not a sizing tool. Real deployments also depend on precision and quantization, context length, whether you're training or only running (inference) the model, and the serving software's own overhead - this only applies the same simplified formula used throughout this store (parameters x 1 GB x 20% overhead), not a substitute for the guided Startup/Mid-size paths or a real quote.
Frequently Asked Questions
Straight answers to the questions this store exists to teach.
What is GPU memory, and why does it decide what I can buy?
Parameters are the numbers a model learned during training - roughly, its "knowledge dials." More parameters usually means a more capable model, but also means more GPU memory is needed just to hold it in memory before it can answer a single question. That memory requirement is what determines which hardware tier you actually need.
Why do you add 20% on top of the raw parameter count?
Required GPU memory = parameters (in billions) x 1 GB x 1.2 (20% working-memory overhead). This is a simplified teaching formula, not a production sizing tool - real-world deployments are also affected by quantization, the speed of the connections between GPUs, and the software framework serving the model. We show the simple version so the relationship between parameters and memory stays visible and checkable by hand.
Are these prices real?
NVIDIA does not publish official list prices for its datacenter GPUs, servers, or racks - these are sold through partners with negotiated pricing. Every price for those products in this store is a market estimate, built from reseller quotes and industry reporting, not a fixed number. Where that applies, the product's card says so.
What's a cluster, and why can't I just buy several gaming PCs instead?
A cluster is a group of machines connected by very fast, dedicated networking so they can work together on one job as if they were a single, much bigger machine. A cluster's power comes from the connections between machines, not just the machines themselves - ordinary desktops can't share GPU memory the way a real cluster can, so a model too big for one PC simply won't run no matter how many desktop PCs you add.
Why don't you sell Llama models?
This store only stocks models genuinely licensed MIT or Apache 2.0 - fully open, with no usage restrictions. Meta's latest large Llama models use a custom license, not MIT or Apache, which requires a separate license from Meta once a product exceeds 700 million monthly active users. Because it doesn't meet the license bar, we don't stock it.
Do I need to create an account or pay to get a quote?
No payment, no account - just tell us who you are and which build you want, and we'll save your request. You'll get a request number, and every request is listed publicly at /requests.