Every year since 2022 someone has declared the used 3090 obsolete, and every year the used 3090 ignores them. The reason is structural, not sentimental: local model serving is gated by memory capacity, not compute. Weights have to fit in VRAM. Once they do, Ampere's tensor cores — five years old now — are plenty fast for a single user reading tokens at reading speed.

Why 24 GB keeps mattering more than speed

Quantization changed the arithmetic. Modern 4-bit quantized (GGUF/AWQ-style) builds of open-weight models cut memory footprints roughly four-fold with modest quality loss, and 2025–2026's model generation got dramatically better at small sizes. The practical map on one 24 GB card, as of late 2026:

  • 7B–14B models run fast with lots of context headroom. Great daily drivers for drafting, summarizing and RAG over personal documents.
  • 32B-class models fit in 4-bit with room for a real context window. This is the sweet spot most home-lab users actually live in.
  • 70B-class models need ~40 GB in 4-bit — that's two 3090s at roughly $1,600 combined, still cheaper than a single current-generation 48 GB card.

And the card doesn't sleep when the LLM doesn't run: Stable Diffusion-class image generation, Whisper transcription, local TTS and vision models all fit comfortably. A 3090 box is a general-purpose local-AI appliance, not a one-trick pony.

The 2026 used-market reality

Three things changed since the "don't buy ex-miner cards" era, and one thing didn't. Datacenter refresh cycles and rental fleets pushed a steady stream of 3090s onto the secondary market, bottoming used prices near $700–$900. The mining-era bogeyman is mostly priced in — these cards have already survived their sweatshop years. And quantized-model quality means a 2022 card runs a 2026 experience. What didn't change: you're buying a used GPU from a stranger. Buy from sellers who show VRAM temperatures under load, not just FPS screenshots. GDDR6X on the 3090 runs hot by design; a card that throttles on memory junction temp in the seller's video will throttle in yours.

Power, noise and the domestic treaty

Plan for ~350 W peak, a PC that idles around 40–60 W, and a machine that is audible under load. At the US average electricity price, 8 hours of daily use costs roughly $15/month — see our TCO spreadsheet for the full line items. Undervolting and power-limit curves typically shave 15–25% of that at single-digit performance cost, and turn the fan curve from hair-dryer to acceptable roommate.

When to buy something else

  • You need 70B+ in one box, quietly. Two 3090s is the budget answer, but it's two loud hot cards; a single used A6000-class 48 GB card or a current 32 GB consumer part may fit your desk and your marriage better.
  • You need maximum tokens/second for many users. A 3090 serves one or two humans gloriously and a team badly. That's a rental job.
  • You need reliability, not experiments. A used consumer card has no ECC, no warranty and a history. If the box must be up, rent — the enterprise math lives in our renting-vs-buying piece.
  • You're training, not serving. Fine-tunes beyond LoRA on a 3090 are an exercise in swap-file patience. Rent H100 hours by the job.

The 3090 isn't exciting in 2026. That's exactly why it keeps winning: it's the last card whose price fell to its floor while its memory still clears the bar. Obsolescence for this one isn't a date — it's a model size. When 70B quantizes into 20 GB and runs on a 16 GB card at full speed, the used 3090 finally retires. Based on the current pace, don't hold your breath this quarter.