A builder's guide to capping NVIDIA cards for Ollama
NVIDIA power limits for local LLMs: we swept 100 to 200 W so you can cap with confidence
We swept an RTX 3060 Ti from 100 to 200 W under Ollama. 150 W cost 3% speed for 25% less draw. Here's the cap we'd pick, the nvidia-smi commands, and how to reapply it at boot.
Cap an RTX 3060 Ti at 150 W: lose 3% speed, save 25% of the watts.
Read your card's rangeAsk nvidia-smi for the Min, Max and Default Power Limit
Set and verify the capnvidia-smi -pl 150, then check the SW Power Cap counter
Reapply at every bootThe limit resets to default after driver unload
Illustrated summary · not a benchmark chart
The short answer
Cap an RTX 3060 Ti at 150 W: lose 3% speed, save 25% of the watts.
On an RTX 3060 Ti running a 7B model in Ollama, cap at 150 W: speed drops 3% and draw drops 25%. Go to 125 W if a cooler card matters more than the last few tok/s; it runs cooler, and likely quieter. Avoid 100 W unless you have to, because it gives up 24% of speed and returns no better tokens per watt than 125 W.
Anyone running 7B–8B models in Ollama on an always-on NVIDIA card who wants lower draw, less heat and likely less fan noise without a meaningful loss in tok/s.
Anyone whose card has a high minimum limit, such as the RTX 2060 at 125 W, and anyone running workloads we didn't test: big contexts, larger models or training.
Our assessment is based on vendor documentation, not a scored hands-on test. How we evaluate products.
We capped an RTX 3060 Ti at five power limits and ran the same Ollama prompt at each one. Our position: most 7B inference rigs are paying for watts they don’t use. The sweep, the commands and the reboot fix are below.
Key takeaway: At 150 W our 3060 Ti produced 75.5 tok/s against 77.8 tok/s at stock 200 W. That is 3% less speed for 25% less draw (bench notes).
Pick 150 W and you keep 97% of the speed
The first 50 W you take off the card cost almost nothing. Our method was qwen2.5:7b Q4_K_M on Ollama 0.34.2, with three 512-token runs per cap at seed 1 and temperature 0, a warm-up first, and driver telemetry logged every 250 ms (bench notes).
| Cap | tok/s | Avg draw | tok/W | Core clock | Peak temp |
|---|---|---|---|---|---|
| 100 W | 58.8 | 99.6 W | 0.59 | 1,092 MHz | 50 C |
| 125 W | 72.3 | 122.5 W | 0.59 | 1,516 MHz | 53 C |
| 150 W | 75.5 | 145.4 W | 0.52 | 1,675 MHz | 57 C |
| 175 W | 76.9 | 172.0 W | 0.45 | 1,789 MHz | 61 C |
| 200 W (stock) | 77.8 | 193.4 W | 0.40 | 1,855 MHz | 66 C |
The memory clock held at 6,801 MHz at every cap, and only the core clock moved. Our read: token generation on a 7B model is memory-bound, so an untouched memory clock is why the curve stays flat down to 125 W. An earlier run the same day at 150 W measured 76.3 tok/s on qwen2.5:7b and 73.9 tok/s on llama3.1:8b, so a second 8B-class model landed in the same range (bench notes).
Drop to 125 W when a cooler card matters more than the last 7%
We consider 125 W the efficiency floor. It matched 100 W at 0.59 tokens per watt and ran 13.5 tok/s faster. Against stock, it gave up 7% of speed for 37% less draw, and peak temperature fell from 66 C to 53 C (bench notes).
Going below 125 W is a bad trade on this card. At 100 W you lose 24% of speed for 49% less draw, and efficiency does not improve at all.
Example (illustrative): Say a batch job generates 100,000 tokens. At stock it takes about 21 minutes at 193 W. At 150 W it takes about 22 minutes at 145 W. That works out to roughly 0.069 kWh versus 0.053 kWh, calculated from our measured tok/s and average draw.
Set the cap with nvidia-smi and confirm it bites
Ask nvidia-smi for your card’s Min, Max, Default and Enforced Power Limit before you type a number. NVIDIA’s docs say -pl sets the maximum power limit in watts, requires root, and must fall between the Min and Max Power Limit that nvidia-smi reports (nvidia-smi docs). Our 3060 Ti reports 100–220 W. An RTX 2060 on our test bench reports only 125–170 W, so a 100 W cap is not universal (bench notes).
sudo nvidia-smi -pl 150
nvidia-smi --query-gpu=clocks_event_reasons_counters.sw_power_cap,clocks_event_reasons_counters.sw_thermal_slowdown
“SW Power Cap” is the reason code for the power algorithm pulling clocks down to hold the limit. Seeing it during a run means the cap is working. Rising thermal-slowdown counters mean cooling, not the cap, is setting your speed (nvidia-smi docs).
Reapply the cap at every boot or you lose it
A cap set by hand does not survive a reboot. NVIDIA documents that the power limit returns to the default after driver unload. Persistence mode (-pm 1) does not persist across reboots either (nvidia-smi docs).
Watch out: If you size your power budget, UPS runtime or electricity estimate around a 150 W card, check after every reboot and driver update. Otherwise you may be running the full 200 W without knowing it.
Our proposed fix is a root-level startup task that runs nvidia-smi -pm 1 and then nvidia-smi -pl 150. A oneshot systemd unit or your distro’s equivalent will do. We have not verified this across driver updates, so check Enforced Power Limit after the first reboot.
Run this checklist before you commit to a cap
- Ask nvidia-smi for the Min, Max and Default Power Limit for your card and note them.
- Warm up the model, then run the same prompt three times at stock, with a fixed seed and temperature 0.
- Set 150 W, repeat the runs, and compare tok/s and average draw.
- Try 125 W. Keep it if the speed loss is acceptable to you.
- Check the
sw_power_capand thermal counters to confirm what is limiting clocks. - Add the boot-time reapply, reboot once, and confirm the Enforced Power Limit.
What we couldn’t verify: we tested one card and one model family at 512-token outputs. Runs went from lowest to highest cap back to back, so later steps started warmer. We did not measure wall power, noise, prompt-processing speed on long contexts, or larger models that spill out of 8 GB.
What we would do next
Three moves before you commit.
Baseline at stock
Warm the model up, then do three fixed-seed runs at temperature 0. Log tok/s and the average draw from nvidia-smi so you have your own card's curve to compare against.
Try 150 W, then 125 W
Set each cap with sudo nvidia-smi -pl, repeat the runs, and check the sw_power_cap and thermal counters. Keep the lowest cap whose speed you can live with.
Make it stick
Add a root startup task that enables persistence mode and sets the limit. Reboot once and confirm the Enforced Power Limit before trusting your power math.
Common questions
Quick answers.
Does a lower power limit slow Ollama down?
Barely, down to about 150 W. On our RTX 3060 Ti, 150 W ran 3% slower than stock and 125 W ran 7% slower. The memory clock stayed at 6,801 MHz at every cap, and only the core clock dropped.
Why does my power limit disappear after a reboot?
NVIDIA's docs say the power limit returns to the default after driver unload, and persistence mode doesn't persist across reboots either. Both have to be reapplied at boot by a root task.
Can I set any wattage I want?
No. The value must fall between the Min and Max Power Limit that nvidia-smi reports. Our 3060 Ti allows 100–220 W. An RTX 2060 on our bench allows only 125–170 W.
How do I know the cap is actually working?
Query clocks_event_reasons_counters.sw_power_cap during a run. If it climbs, the power algorithm is pulling clocks down to hold your limit. If the thermal slowdown counters climb instead, cooling is what's limiting you.
Sources & methodology
Where these facts come from.
- racksandwatts.com/lab/bench-notes-2026-09/ · checked September 25, 2026
- docs.nvidia.com/deploy/nvidia-smi/index.html · checked September 25, 2026
- www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3060-3060ti/ · checked September 25, 2026
We read the manufacturer's specifications and the retailer listings on the dates shown. We have not run a hands-on test unless the article says so.
Ready to look for yourself?
Open NVIDIA's product page with this guide beside it.
Use the worksheet above while you are on their site. If the VRAM or the watts do not fit your plan, you will know before you buy.
See NVIDIA's current pricing ↗Prepared by Racks & Watts. Measurements are from our own hardware unless the article says otherwise; prices change, verify on the retailer's page before you buy.