Racks & Watts
We may earn a commission when you buy through links on this site, at no extra cost to you. Our recommendations are based on genuine research, not commission. Learn more.

A builder's guide to budget GPUs for Ollama

RTX 3060 Ti vs RTX 2060 for Ollama: we benchmarked each card alone

We benched an RTX 3060 Ti and an RTX 2060 alone in Ollama. The 2060 runs 7B models at 76% speed, drops to 61% on llama3.1:8b, and neither card fits a 14B model.

NVIDIAHead-to-head

The 2060 holds up on 7B models, then falls off a cliff at 8B.

Size your model first7B fits both cards; llama3.1:8b spills off the 2060

Cap the power150 W on the 3060 Ti, 125 W on the 2060: ~97% of stock speed

Price the premium~$127 more used buys 64% more speed on llama3.1:8b

Illustrated summary · not a benchmark chart

qwen2.5:7b speed75.8 vs 57.7 tok/s
llama3.1:8b speed73.6 vs 44.8 tok/s
Idle board power16.5 W vs 6.9 W
Research checkedSeptember 26, 2026

The short answer

The 2060 holds up on 7B models, then falls off a cliff at 8B.

Keep a spare RTX 2060 for 7B models at a 125 W cap. If you're buying for 8B models like llama3.1:8b, pay the roughly $127 used-price premium for the 3060 Ti. If you need 14B, neither card alone is enough.

A stronger fit

The 3060 Ti fits builders running 8B models all day who want every layer on the GPU. The 2060 fits anyone who already owns one and sticks to 7B-class models like qwen2.5:7b.

Think twice if

Skip both if 14B models are the goal: neither card alone kept qwen2.5:14b in VRAM. Skip the 3060 Ti if the box idles 24/7 and 7B speed is enough, since it idles at 2.4x the 2060's board power.

Our assessment is based on vendor documentation, not a scored hands-on test. How we evaluate products.

Plans at a glance

NVIDIA's tiers, side by side.

RTX 3060 Ti new

$395.00 new
8 GB GDDR6, 256-bit, 200 W

For builders who want a new card with every layer of an 8B model on the GPU.

Check current terms ↗

RTX 3060 Ti Renewed

$299.99 Renewed
Dell OEM listing, 8 GB

The cheapest listed 3060 Ti. It's our pick if you accept a refurbished card.

Check current terms ↗

RTX 2060 used

From $185.99 used
6 GB GDDR6, 192-bit, 170 W

For 7B-only workloads on a tight budget.

Check current terms ↗

RTX 2060 new

$259.00 new
Zotac compact, 6 GB

Hard to justify: a Renewed 3060 Ti costs about $41 more.

Check current terms ↗

Prices as listed on the retailer or manufacturer page on September 26, 2026. Listings and stock change daily; verify before you buy.

Key takeaway: On qwen2.5:7b, the RTX 2060 delivers 76% of the RTX 3060 Ti’s speed. On llama3.1:8b that drops to 61%, because the model no longer fits in the 2060’s VRAM. Pick the card by model size, not by the spec sheet.

Know exactly where the 2060 stops being “good enough”

The 2060 is fine at 7B and falls apart the moment a model spills out of VRAM. We ran each card alone on our test bench, each in its own Ollama 0.34 instance with nothing else on the GPU. Settings were flash attention on and q8_0 KV cache, with three 512-token runs per row at seed 1 and temperature 0 after a warm-up. Speed is the mean of the 3 runs, draw is the mean of runs 2–3, and temperature is the peak (bench notes).

Model RTX 3060 Ti (150 W cap) RTX 2060 (125 W cap) On GPU (Ti / 2060)
qwen2.5:7b 75.8 tok/s 57.7 tok/s 100% / 100%
llama3.1:8b 73.6 tok/s 44.8 tok/s 100% / 86%
qwen2.5:14b 14.5 tok/s 8.6 tok/s 71% / 48%

The llama3.1:8b row is the one that matters. The default Q4_K_M tag is about 4.7 GB on disk. Runtime VRAM is weights plus KV cache and overhead, so it doesn’t fully fit on the 2060 (Ollama library). With 14% of the model on the CPU, the 2060’s GPU sat at 51–55 W, waiting on the CPU (bench notes).

The specs explain the gap. The 3060 Ti has 8 GB on a 256-bit bus and 4864 CUDA cores (NVIDIA). The standard 2060 has 6 GB on a 192-bit bus and 1920 CUDA cores (NVIDIA).

Cap both cards: stock power buys you almost nothing

Run the 3060 Ti at 150 W and the 2060 at 125 W. You give up about 2–3% of speed.

On qwen2.5:7b, the 3060 Ti went from 75.8 tok/s at 150 W to 77.9 tok/s at 200 W. That’s 97% of stock speed on 144.2 W instead of 191.9 W, and efficiency rose from 0.41 to 0.53 tok/W. The 2060 went from 57.7 tok/s at 125 W to 59.1 tok/s at 170 W. The cap cut its peak temperature from 80 C to 64 C (bench notes).

The software power ranges are 100–220 W on the 3060 Ti and 125–170 W on the 2060, so both caps sit inside what the driver allows (bench notes). On models that spill to the CPU, the cap does nothing. The 2060 ran llama3.1:8b at 44.8 tok/s at 125 W and 45.0 tok/s at 170 W, and both cards ran qwen2.5:14b at the same speed at either cap.

Tip: Set the cap before you shop for a PSU. NVIDIA lists 600 W required system power for a 3060 Ti at its 200 W stock rating (NVIDIA). Our tests ran it at 150 W with a 3% loss on qwen2.5:7b.

Price the premium per token, not per card

The 3060 Ti’s premium pays off for 8B models and barely at all for 7B. Prices below are Amazon listings from September 25–26, 2026 (3060 Ti search, 2060 search):

  • Used: 3060 Ti offers started at $313.22 and 2060 offers at $185.99, a $127.23 gap.
  • New: $395.00 vs $259.00, a $136 gap.
  • Refurbished: A Renewed 3060 Ti listed at $299.99, only $41 more than the cheapest new 2060.

Example (illustrative math from those prices): On qwen2.5:7b at the caps above, a used 3060 Ti works out to 75.8 / $313.22 = 0.24 tok/s per dollar. A used 2060 works out to 57.7 / $185.99 = 0.31, so the 2060 wins on value. On llama3.1:8b the numbers are 0.235 for the Ti vs 0.241 for the 2060, a near tie. The Ti also keeps the whole model on the GPU, so it wins in practice.

Watch out: Don’t buy a new RTX 2060. Listings ran $259–$369.99, and several sit above the $299.99 Renewed 3060 Ti. You’d pay more for a card with less VRAM.

Count idle watts if the box never sleeps

The 2060 is the better idler. With no model loaded, we measured 16.5 W of board power on the 3060 Ti and 6.9 W on the 2060 (bench notes).

That 9.6 W gap adds up to about 230 Wh per day, or roughly 84 kWh per year of 24/7 idling (illustrative arithmetic). Multiply by your own electricity rate. For a box that sits idle most of the day, that recurring cost belongs next to the purchase price.

Use this worksheet to pick your card

  1. List the largest model you’ll run every day, not the one you’ll try once.
  2. If it’s 7B-class (qwen2.5:7b), the 2060 at 125 W gets you 57.7 tok/s fully on the GPU. Keep the spare card or buy used from $185.99.
  3. If it’s llama3.1:8b, only the 3060 Ti kept 100% on the GPU, at 73.6 tok/s. Budget $299.99–$313.22 for a Renewed or used card.
  4. If it’s 14B, neither card alone is enough (14.5 and 8.6 tok/s with partial CPU offload). Earlier the same week, qwen2.5:14b split across both cards fully in VRAM ran at about 38 tok/s. That result came from a separate run, not the solo bench above (bench notes).
  5. Set the power cap on day one: 150 W for the 3060 Ti, 125 W for the 2060.
  6. Add the idle gap (9.6 W) to your running cost if the box stays on 24/7.

What we couldn’t verify: We didn’t bench the 12 GB RTX 2060 variant (2176 CUDA cores), and its extra VRAM could change the 8B result. We didn’t test longer contexts beyond 512-token runs. Used and Renewed card condition and listing prices will vary.

What we would do next

Three moves before you commit.

Check whether your daily model fits in VRAM

Run it with ollama ps and look at the GPU percentage. If a 2060 shows less than 100%, as llama3.1:8b did at 86%, you're CPU-bound and a power cap or tuning won't help.

Cap power before anything else

Set 150 W on a 3060 Ti or 125 W on a 2060. On qwen2.5:7b we lost about 3% of speed, and the 2060's peak temperature fell from 80 C to 64 C.

Buy Renewed or used, not a new 2060

If 8B models are the target, a $299.99 Renewed 3060 Ti beats any new 2060 listing. If 7B is enough, a used 2060 from $185.99 is the value pick.

Common questions

Quick answers.

How much slower is an RTX 2060 than a 3060 Ti in Ollama?

On our test bench with both cards capped, the 2060 ran qwen2.5:7b at 57.7 tok/s vs 75.8 (76%) and llama3.1:8b at 44.8 vs 73.6 (61%). The bigger gap on llama3.1:8b comes from 14% of the model spilling to the CPU on the 2060.

Why doesn't llama3.1:8b fit on the 2060 if it's only 4.7 GB?

The 4.7 GB is the size on disk. Runtime VRAM also needs room for the KV cache and overhead, and on our bench only 86% of the model landed on the 2060's GPU.

Is stock power worth it for LLM inference?

Not on these cards. The 3060 Ti gained 2.1 tok/s going from a 150 W to a 200 W cap on qwen2.5:7b, and the 2060 gained 1.4 tok/s going from 125 W to 170 W. On models that spill to the CPU, the cap changed nothing.

Can either card run a 14B model?

Not well on its own. With partial CPU offload, qwen2.5:14b ran at 14.5 tok/s on the 3060 Ti and 8.6 tok/s on the 2060. Split across both cards fully in VRAM, a separate run earlier the same week hit about 38 tok/s.

Which card idles lower?

The 2060 does. We measured 6.9 W of board power with no model loaded, vs 16.5 W on the 3060 Ti.

Sources & methodology

Where these facts come from.

We read the manufacturer's specifications and the retailer listings on the dates shown. We have not run a hands-on test unless the article says so.

Ready to look for yourself?

Open NVIDIA's product page with this guide beside it.

Use the worksheet above while you are on their site. If the VRAM or the watts do not fit your plan, you will know before you buy.

See NVIDIA's current pricing ↗

Prepared by Racks & Watts. Measurements are from our own hardware unless the article says otherwise; prices change, verify on the retailer's page before you buy.