RECOMPUTING THE PER-CHIP RENTAL PRICE — the exact specification compute.pangle.online/recompute.txt · v0.4 · 2026-10-02 (v0.4: wording only, from an outside reviewer's full recompute on 2026-10-01. No published value changed.) WHY THIS FILE EXISTS /methodology/ describes the method in prose. Prose is not enough to reimplement from: it says "median" in places where the code means different things, and it does not pin rounding, tie-breaks or cut-offs. This file states what the code actually does, precisely enough that an independent implementation can be built from it and compared against ours offer by offer. It describes BEHAVIOUR AS SHIPPED. It is a description, not a defence. If your implementation disagrees with ours and this file says we are right, this file is the thing to attack. SCOPE The per-chip dollar price at /api/chip-index — "what does an hour of an H100 cost". Not the chained basket level at /api/index, which is a different number answering a different question. That one is specified in /recompute-index.txt. INPUTS ⚠️ CORRECTED 2026-09-17 — v0.1 of this file named the wrong file and the wrong column names, which would have stopped a stranger at step zero. Sorry. The tape is /data/days/gpu-price-tape-YYYY-MM-DD.csv.gz, one gzipped CSV per UTC day. (/data/days/index.json is the MANIFEST, ~600 bytes, not the data.) Header, verbatim: snapshot_ts_utc_epoch,provider,gpu_model,offer_class,offer_count, min_price_per_gpu_hr_usd,median_price_per_gpu_hr_usd All prices are USD per GPU per hour. All arithmetic below is IEEE-754 double precision. From 2026-09-17 each sealed day also has a companion file, /data/days/gpu-price-tape-median-true-YYYY-MM-DD.csv.gz, listed with its sha256 in /data/days/median-true.json. Header: snapshot_ts_utc_epoch,provider,gpu_model,offer_class, median_true_price_per_gpu_hr It joins 1:1 to the sealed tape on the first four columns. Companion files are not part of the merkle roots; check each one against the sha256 in median-true.json. STAGE 1 — per provider, per snapshot (written at poll time) For one provider, one gpu_model and one class, take that poll's list of per-GPU prices, sort ascending, and store: min = prices[0] median = prices[len(prices) // 2] ⚠️ THE SEALED COLUMN IS NOT THE STATISTICAL MEDIAN. It is the UPPER-MIDDLE order statistic. For an even-length list it returns the higher of the two middle values, not their mean: [10, 20] gives 20, not 15. The column is named median_price_per_gpu_hr and /methodology/ calls it a median. On even-length lists it is not one, and it biases high. From 2026-09-17 07:50:57Z (snapshot_ts 1789631457) the per-provider price used from Stage 2 on is the true median, statistics.median(prices), published in the companion file; before that it is the upper-middle value as sealed. STAGE 2 — binning Rows are read ordered by snapshot_ts ascending and assigned to 30-minute bins: bin = (snapshot_ts // 1800) * 1800 Within a bin, the LAST row read for a given provider wins, so the latest snapshot in the bin is the one used. The key is the PROVIDER alone, not provider plus gpu_model: in a series of several model strings (h100-nvl), a provider appears in a bin at most once and its last row in tape order wins across all of the series' models. (Measured 2026-10-01: no provider has ever listed both H100 NVL strings in the same bin, so this has not yet mattered.) Rows whose per-provider price (the companion value where the row has one, else median_price_per_gpu_hr_usd) is zero or less are excluded before binning. ON TIES: v0.1 left this open. Measured since on every sealed day, 2026-08-23 to 2026-09-30 (by an outside reviewer, 2026-10-01): 530,439 rows, ZERO duplicate (snapshot_ts, provider, gpu_model, offer_class) groups — the writer emits one row per group per poll, so the tie cannot arise in practice. If you ever see a duplicate group, treat it as a defect in our data and tell us; do not guess a winner. STAGE 3 — the bin's price, across providers bin_median = round(statistics.median(provider_medians), 4) bin_low = round(min(provider_lows), 4) # lows of zero or less excluded; None if none left Here "median" IS the statistical median: for an even number of providers it is the MEAN OF THE TWO MIDDLE VALUES. This is a different rule from Stage 1. The median is across PROVIDERS, never across offers — a provider listing four hundred cards does not outvote one listing a single card. STAGE 4 — the day's price day_median = round(statistics.median(bin_medians_for_that_UTC_day), 4) day_low = round(min(bin_lows), 4) Days are UTC midnight to midnight, the same boundary the sealed tape uses. A bin whose provider lows are all zero or less has bin_low None. Such a bin still counts toward the day's median and toward bins; it is skipped when taking the day's low. If every bin of the day has bin_low None, day_low is None and is published as null. A bin is never dropped for lacking a low. providers_max is the largest number of distinct providers in any one bin of the day after the exclusions; bins is the number of bins with at least one surviving row. ROUNDING Every published price is rounded once at Stage 3 (bin median, bin low) and once at Stage 4 (day median, day low), each time to four decimal places by IEEE-754 round-half-to-even applied to the binary double, which is Python's built-in round(x, 4). This is not decimal half-up rounding: a value whose decimal expansion ends in 5 at the fifth place may round down, because its nearest double lies below the half. Example: the bin median 2.66665 is stored as the double 2.6666499999999997 and rounds to 2.6666, not 2.6667. Implementations in other languages must round the double, not the printed decimal. Stage 4 takes the median of the already-rounded Stage 3 values. (Wording from an outside implementer, adopted 2026-10-01.) WHAT IS EXCLUDED FROM THE HEADLINE - class: the headline is on_demand only. interruptible is published beside it and never merged into it. - form factor: H100 SXM, PCIE and NVL are separate series. A provider stating only "H100" with no form factor goes into its own ambiguous series and is excluded from the headline rather than guessed into one of the others. SERIES TABLE (normative, versioned with this file) A series is a fixed set of gpu_model strings; rows whose gpu_model is in no series are ignored. Matching is exact and case-sensitive on the tape string. h100-sxm = {H100 SXM} h100-pcie = {H100 PCIE} h100-nvl = {H100 NVL, H100-NVL-PCIE} h100-unspecified = {H100} b200 = {B200} b200-cc = {B200 CC} Headline series: h100-sxm, h100-pcie, b200. SYNTHETIC TEST VECTORS (written by an outside implementer, 2026-10-01) One synthetic UTC day, 2026-10-01, series X100 = {X100}. Tape rows in header order (snapshot_ts_utc_epoch, provider, gpu_model, offer_class, offer_count, min, median): 1790812860,alpha,X100,on_demand,3,2.0,2.5 1790812900,alpha,X100,interruptible,3,0.8,1.2 1790812920,beta,X100,on_demand,2,2.1,2.7 1790813700,gamma,X100,on_demand,5,1.9,2.9 1790814500,alpha,X100,on_demand,3,2.05,2.6 1790814600,alpha,X100,on_demand,3,2.05,2.6 1790815200,gamma,X100,on_demand,5,1.95,2.9 1790816500,alpha,X100,on_demand,3,0.0,2.5333 1790816600,beta,X100,on_demand,2,-1.0,2.8 1790818201,delta,X100,on_demand,1,3.1,3.1 1790818300,beta,X100,on_demand,2,2.0,0.0 1790899200,alpha,X100,on_demand,3,1.0,1.0 Expected, on_demand: median 2.725, low 1.9, providers_max 3, bins 4. Bin medians 2.7, 2.75, 2.6666, 3.1; bin lows 1.9, 1.95, null, 3.1. Expected, interruptible: median 1.2, low 0.8, providers_max 1, bins 1. The last row is the next UTC day and must not appear in 2026-10-01. These exercise: a provider's later row replacing its earlier one in a bin; a row exactly on the 1800-second boundary; an even provider count in a bin; a bin whose lows are all zero or less; a median of 0.0 dropped before binning; the double-rounding example above; an even bin count at Stage 4. Our production code (core/chip_index.py) gives exactly these values. REFERENCE TEST VECTORS /vectors.json carries the expected output for one real sealed day (2026-09-16) for every chip and both classes. Download that day's tape from /data/days/, run your implementation, and compare. If you match, you have the method. If you do not, one of us is wrong and we would like to know which. ⚠️ A LIMIT ON "REPRODUCIBLE BY A STRANGER", stated because it is easy to miss. The vectors — and the sealed tape itself — begin at STAGE 2. The tape's median_price_per_gpu_hr column is already the OUTPUT of Stage 1, so Stage 1 is NOT independently checkable from anything we publish: the raw per-offer price lists a provider returned at poll time are never published, only their summary. An outsider can verify Stages 2-4 exactly. Stage 1 is not reproducible from published data — but it is NOT invisible either, and v0.1 of this file overstated that. Both min_price_per_gpu_hr_usd and median_price_per_gpu_hr_usd are published per row, so wherever offer_count is 2 the whole price list is recoverable: the stored "median" is the MAX and the true median is the mean of the two. You can measure the Stage 1 bias yourself on exactly those rows, and you should. THE STAGE 1 BIAS, STATED PLAINLY The upper-middle rule (the sealed column, and the price used before 2026-09-17 07:50:57Z) always reads high on even-length offer lists. HOW BIG IS IT. Measured on the 2026-09-16 tape, correcting ONLY the rows where offer_count is 2 (the ones a stranger can check): H100 SXM 3.49 -> 3.39, +2.9%. B200 6.79 -> 6.385, +6.3%. H100 PCIE unchanged. That is a FLOOR, not the full number: it corrects only n=2, and 40.9% of on-demand rows that day were even-length. Recomputed against the raw price lists the figure for H100 SXM is nearer +10%. The bias is upward in every case, because the rule always takes the higher of the two middle values. If you are reimplementing: for each sealed row, use the companion file's true median where it has one and the sealed median_price_per_gpu_hr_usd otherwise.