Learn · audit

Tuning

How Ember reaches a good operating point without third-party software, every figure a dated measured row or marked approximate.

Tuning: how Ember reaches a good operating point, and what the kit builds with

8 October 2026. The register row (plan p. 14): "Published optimisation work, compiler settings and safe tuning logic: Ember reaches good operating points without third-party software." Served under igneum.network/miner. Every figure here is a measured row of the day it names, or marked approximate.

1. The knee rule

A card's hash rate on the Igneum hash is bound by dependent memory reads, not by the core clock. The core clock can fall a long way before the rate moves, and the card's power falls with it. The knee is the lowest core clock that holds the rate. Ember locks the card at the knee. Nothing else is touched: no voltage, no fan curve, no memory overclock.

CardKnee (SM MHz)Rate at the kneeWatts at the kneeRate at stockWatts at stockEnergy per hash, knee against stockMeasured
RTX 50901,300132.7 MH/s314.8 W140.0 MH/s456.1 W2.37 against 3.26 µJ, minus 27 percent8 Oct 2026, class v5 genesis pack, the project’s own rig, nvidia-smi 1 Hz
RTX 50801,10060.3 MH/s123 Wapproximate 61 MH/sapproximate 200 Wapproximate minus 38 percent8 Oct 2026, Ember climb, the project’s own rig (the stock row approximate)
RX 9070 XTno core lock lever in 0.3.20; the AMD knob is a clock offset and a power limit19.0 MH/s at minus 500 MHz and minus 30 percent149.3 W18.2 MH/sapproximate 300 W0.127 MH per W at the knob against approximate 0.068 Oct 2026, the AMD grid, the project’s own rig
RX 7600the same knob; grid owed (the card must be mining for the grid)13.9 MH/s113 Wthe samethe same0.123 MH per W at stock8 Oct 2026, card-in, the project’s own rig

The rule, as the app applies it: on NVIDIA the core clock is locked at the knee with the driver's own lock (nvidia-smi -lgc 0,<knee>), through the Igneum Power Helper task so the app never runs elevated; on AMD the knob is the clock offset and the power limit through the vendor's own interface; on Apple there is no lever and the card runs at stock. A lock is released (-rgc) when the app stops mining, when a job ends, and on every start.

2. The ladders and the priors

Ember does not start from zero. The search has a prior per architecture and walks from it; the measured knee decides.

ArchitecturePrior kneeWhere it came from
Blackwell, RTX 50901,300 MHzthe efficiency passes of 7 and 8 October 2026
Blackwell, RTX 50801,000 MHzthe same
Ada (RTX 40)2,400 MHznine capped-against-uncapped pairs on rented cards, class v4 (a power cap holds the full rate until the SM clock falls under about 2,400)
Ampere (RTX 30, A-series)1,800 MHzthe same pairs (a 3080 Ti capped to 63 percent with the SM at 749 MHz loses 24 to 28 percent)
anything elsenonethe full ladder

The ladders: the power ladder's rungs are 100, 90, 80, 70, 60 and 50 percent of the card's limit; the clock ladder's rungs are 100, 90, 80, 70, 60, 50 and 45 percent of the card's maximum. The power ladder starts at the rung the prior clock implies (the draw follows the clock cubed on the voltage and frequency curve, rounded up to the next rung). Ember 2 adds a hill-climb from the start point: each probe moves the memory clock up by one step or the core clock down by one step; a probe that produces a refused row (an invalid hash, a fault, a rate under the floor) backs that knob off and is never retried in the run.

Every row of a search carries the clock, the power percent, the memory clock, the measured watts, the rate and the rate per watt. Three tiers come out of the rows:

TierRule
efficiencythe usable row with the most MH per W
balancedthe row within 1 percent of the best rate with the fewest watts
maxthe usable row with the highest rate

A card with no lever has one tier, its stock row, and the note says why. A tier is remeasured when the class changes (the rows are keyed to the class the search ran on) and on the app's own schedule.

3. Safe defaults the app applies per class

ClassWhat the app does on a fresh cardWhy
v3 (the devnet's class until the class v6 cut)the prior knee as the clock cap, the power ladder from its implied rungthe knee rule above
v4 and v5 (the shadow block, the state leaves)the same priors; the v5 knee on the 5090 reads the v4 rate at +2 percent watts (8 Oct 2026)the shadow block runs in the memory shadow; the state leaves cost nothing on the card
v6 (the index fold, the re-weight table, the 64-register window)the same priors; the window halves the hash rate at about 7 percent less card power (a window hash is twice the work) and the tiers are remeasured on the class flipthe register budget changes (section 5)

The app never writes a setting it cannot restore. The lock is released on stop and on start. A tuning failure leaves the card at stock. A tuning file (--tuning <file>, or IGNEUM_TUNING_FILE) carries the chosen worker variant per card so the race (section 4) does not rerun on every start.

4. The worker's flags and what each costs

The CUDA worker (igneum-worker-cuda) compiles each pack's kernel at run time with NVRTC. The OpenCL worker (igneum-worker-opencl) builds it with the vendor's OpenCL compiler. The Metal worker builds it with Apple's. The text of the kernel is the pack's; the worker adds nothing to the hash.

FlagMeaningCost or effect
--device Dthe cardnone
--batch-log2 Bnonces per dispatch, 2^B (default 22)more nonces per dispatch amortise the launch; 24 is the bench setting
--block-warps Wwarps per thread block (default 1)the 5090 and 4090 read within 3 percent across 1, 2, 4 and 8 on every class measured today (8 Oct 2026)
--arch sm_XY, compute_XY, autothe NVRTC target (default the device's own)a PTX target is JIT-compiled by the driver once per pack
--race on, off, a,b,crace the variants of section 5 on first use, or none, or a listone race per pack, about 2 seconds per variant, kept in the tuning file
--variant <name>one variant, no racenone
--check --pack <dir>compile, build the cache and dataset, run the self-test, print timings, exit 0 or 1the gate every pack passes before it serves a job
--bench --pack <dir> --batches Ntime N dispatches, print the fingerprint of the 2^B outputs at base nonce 0the measurement every row on this site comes from
--memprobethe card's dependent and independent read latencies at 4, 64 and 1024 MiBa diagnostic, no hash

The self-test, run before any job: the cache head, last line and FNV-1a 64; the dataset head, last word and 64 samples; the three vector warps of the pack through the bound kernel. A pack that fails is refused.

5. Compiler settings the kit builds with

BackendCompiler and optionsRegister budgetNotes
CUDA (NVRTC in the worker)the device's sm_XY, --std=c++17, -default-device; a variant may add --maxrregcount=N or __launch_bounds__class v5: 48 registers, 24 blocks per SM on the 5090; the window class: 88 to 96 registers on the 5090 (20 blocks per SM), 87 to 104 on the 4090 (16 to 20), no spill (8 Oct 2026, ptxas)no fast-math, no unsafe flag: the hash is integer only
CUDA (the offline bench, proto-cuda/build.sh)nvcc -O3 -std=c++17 -arch=sm_120 (or native)the sameneeds CUDA 12.8 or newer
OpenCLclBuildProgram with -cl-std=CL1.2, CL2.0 or CL3.0 by the device's version and exchange mode; --build-opts appendsthe vendor'sAMD reports the card by its gfx name (the RX 7600 is gfx1102, the RX 9070 XT gfx1201)
MetalApple's compiler, the pack's .metal texts, no optionsthe compiler'sthe M5 Max and the Mac mini M6 run at stock

The variants the worker races, by name: base (the pack's text, one warp per block), w2, w4, w8 (warps per block), u2, u8 (loop unroll), ldg, ldcg, ldcs (the load path), r32, r64 (a register cap), lb4-w4, lb8-w2 (launch bounds), and the pairs u2-ldg, u2-w4, ldg-w4, ldcg-w4. On the 5090 and 4090 the base variant wins or ties on every class measured on 8 October 2026; the race exists for cards the team does not own.

6. What a miner can check

Every row above is reproducible with the kit's worker and the public packs: the fingerprint printed by --bench is the same on every backend and every card (the class v6 all-together pack reads 59e6708e46f1e87c on CPU, CUDA, Metal, Apple OpenCL, an RTX 5090 and an RX 7600; 8 October 2026). A different fingerprint is a bug, and the card's rate is not a figure of merit until it matches.

Nothing here needs third-party software: the lock is the driver's own, the knob is the vendor's own, the worker is the kit's.

Generated from docs/build/tuning.md in the repository at build time. Times are UTC.