Learn AI Series (#134) - AI Infrastructure Economics

Words
4324
Reading
20 min
Listen
Play
3M

Learn AI Series (#134) - AI Infrastructure Economics

variant-a-02-gold.png

What will I learn

  • What GPU compute actually costs across the big cloud providers -- and why the same A100 can cost you three times more depending on where you rent it;
  • spot and preemptible instances, the single biggest lever you have: how to knock 60-90% off a training run without changing a line of your model code;
  • the on-premise versus cloud crossover, and the one number (utilization) that decides it for you;
  • cost-optimization tricks that shrink the bill without touching model quality;
  • the make-versus-buy decision -- train your own model or just call an API -- reduced to a calculation you can actually run;
  • and how to budget an AI project across its whole life so you do not spend the entire war chest on training and have nothing left to serve the thing ;-)

Requirements

  • A working modern computer running macOS, Windows or Ubuntu;
  • An installed Python 3(.10+) distribution with PyTorch (pip install torch); everything here runs fine CPU-only, since the code is about reasoning about compute, not consuming it;
  • You've been through episode #60 (Training LLMs), #125 (GPU Programming Basics) and #126 (Distributed Training) -- and last episode's #133 on synthetic data, whose exercises we settle right below.

Difficulty

  • Beginner

Curriculum (of the Learn AI Series):

Learn AI Series (#134) - AI Infrastructure Economics

I ended episode #133 with a deliberately uncomfortable question. We had spent the whole thing learning to manufacture data -- spinning up rows and images out of thin air with CTGAN and TVAE -- and then I pointed at the elephant nobody in the tutorials likes to name: all of that assumed we had the GPUs to train the generators, the storage to hold both datasets, and the budget to run 300 epochs without flinching. Generating a million synthetic rows is cheap. Serving a model to a million users is emphatically not.

So today we talk about money. Real money -- GPUs, dollars, watts, cloud invoices. This is the episode nobody teaches, and it is the one that decides whether your project ships or dies. Not the algorithm, not the data, not the talent. The money. Here we go ;-)

Nota bene: this episode has less code than usual, and that is on purpose. Infrastructure economics is fundamentally about decisions you make before the first line of training code. The code we do write is code that helps you reason about compute, not code that consumes it. Get this stuff wrong and you burn cash you did not have to burn. Get it right and you can do serious AI work on a budget that does not require a venture capitalist on speed-dial.

Solutions to episode #133's exercises

Homework first, as always -- full code, no hand-waving.

Exercise 1 -- watch SMOTE hallucinate. The task was to build a tiny curved (crescent) minority class, oversample it with SMOTE, and see where SMOTE drops points that could never occur in the real crescent.

import numpy as np
import matplotlib.pyplot as plt
from sklearn.datasets import make_moons
from imblearn.over_sampling import SMOTE          # pip install imbalanced-learn

# make_moons gives us two interleaving crescents; we make ONE of them rare
X_moons, y_moons = make_moons(n_samples=400, noise=0.05, random_state=0)
minority = X_moons[y_moons == 0][:40]              # 40 real crescent points (rare)
majority = X_moons[y_moons == 1]                   # ~200 majority points
X = np.vstack([minority, majority])
y = np.hstack([np.zeros(len(minority)), np.ones(len(majority))])

X_res, y_res = SMOTE(random_state=0, k_neighbors=5).fit_resample(X, y)
synthetic = X_res[len(X):]                          # everything SMOTE invented

plt.scatter(majority[:, 0], majority[:, 1], s=8,  c="lightgray", label="majority")
plt.scatter(minority[:, 0], minority[:, 1], s=25, c="navy",     label="real minority")
plt.scatter(synthetic[:, 0], synthetic[:, 1], s=15, c="crimson", marker="x", label="SMOTE")
plt.legend(); plt.title("SMOTE on a curved manifold"); plt.tight_layout(); plt.show()

The two sentences: SMOTE draws a straight line between a real minority point and one of its neighbours, then drops a synthetic point somewhere on that chord. Where the crescent curves, those chords cut straight across the empty concave belly of the moon -- so SMOTE cheerfully plants "real fraud" in a region no genuine crescent point ever visits. That is the whole "linear interpolation cannot follow a curved manifold" claim, made visible in one scatter plot.

Exercise 2 -- run the three pillars for real. Fit a synthesizer on a small tabular set, sample a synthetic copy, and score it on utility, fidelity and privacy.

import numpy as np, pandas as pd
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score
from sklearn.neighbors import NearestNeighbors

df = load_breast_cancer(as_frame=True).frame       # features + 'target'
train, test = train_test_split(df, test_size=0.3, random_state=42, stratify=df["target"])

# Fit a synthesizer. sdv's CTGAN if available; a per-class gaussian if not.
try:
    from sdv.single_table import CTGANSynthesizer
    from sdv.metadata import SingleTableMetadata
    meta = SingleTableMetadata(); meta.detect_from_dataframe(train)
    synth = CTGANSynthesizer(meta, epochs=100); synth.fit(train)
    fake = synth.sample(len(train))
except Exception as e:
    print("sdv not installed, using a gaussian fallback:", e)
    rng, parts = np.random.default_rng(0), []
    for label, grp in train.groupby("target"):
        cols = grp.drop(columns="target")
        s = rng.normal(cols.mean().values, cols.std().values + 1e-9, size=cols.shape)
        f = pd.DataFrame(s, columns=cols.columns); f["target"] = label
        parts.append(f)
    fake = pd.concat(parts, ignore_index=True)

feat = [c for c in df.columns if c != "target"]

# Pillar 1 -- UTILITY: train on synthetic, judge on REAL held-out data
def auc_from(frame):
    m = RandomForestClassifier(n_estimators=200, random_state=0).fit(frame[feat], frame["target"])
    return roc_auc_score(test["target"], m.predict_proba(test[feat])[:, 1])
print(f"utility  real-train AUC:      {auc_from(train):.3f}")
print(f"utility  synthetic-train AUC: {auc_from(fake):.3f}")

# Pillar 2 -- FIDELITY: how far apart are the correlation structures?
corr_gap = (train[feat].corr() - fake[feat].corr()).abs().mean().mean()
print(f"fidelity mean |corr diff|:    {corr_gap:.3f}  (lower is better)")

# Pillar 3 -- PRIVACY: nearest REAL neighbour of each synthetic row
nn = NearestNeighbors(n_neighbors=1).fit(train[feat].values)
dmin = nn.kneighbors(fake[feat].values)[0].ravel()
print(f"privacy  5th-pct dist to real: {np.percentile(dmin, 5):.3f}  (near 0 = memorised)")

Which pillar looked weakest? On this dataset, with the gaussian fallback, fidelity is the loser -- the per-class gaussian throws away every correlation between columns, so the correlation gap is embarrassingly large even though the utility AUC stays respectable. That is the whole lesson in one run: a synthetic set can score fine on the task you tested and still be a bad model of reality. Real CTGAN closes the fidelity gap a lot, at the cost of training time -- which, funnily enough, is exactly what this episode is about.

Exercise 3 -- break it on purpose. Keep ~30 real positives, blow them up into 3,000 synthetic positives, and prove that "more rows" did not buy you "more diversity".

import numpy as np
from sklearn.datasets import make_blobs
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import recall_score

rng = np.random.default_rng(0)
# Positives really live in TWO sub-clusters (A and B). Our 30 samples only see A.
posA, _ = make_blobs(n_samples=200, centers=[[0, 0]], cluster_std=0.5, random_state=1)
posB, _ = make_blobs(n_samples=200, centers=[[6, 6]], cluster_std=0.5, random_state=2)
neg,  _ = make_blobs(n_samples=2000, centers=[[3, -3]], cluster_std=1.5, random_state=3)

seen = posA[:30]                                   # the only 30 real positives we "have"
idx = rng.integers(0, 30, size=3000)               # naive synthetic: jittered copies
synth_pos = seen[idx] + rng.normal(0, 0.1, size=(3000, 2))

X_train = np.vstack([synth_pos, neg[:1000]])
y_train = np.hstack([np.ones(3000), np.zeros(1000)])

# Held-out REAL test: positives from BOTH clusters, including the unseen B
X_test = np.vstack([posA[30:130], posB[:100], neg[1000:1200]])
y_test = np.hstack([np.ones(200), np.zeros(200)])

clf = RandomForestClassifier(n_estimators=200, random_state=0).fit(X_train, y_train)
pred = clf.predict(X_test)
print(f"recall on cluster A (seen):   {recall_score(y_test[:100],   pred[:100]):.2f}")
print(f"recall on cluster B (unseen): {recall_score(y_test[100:200], pred[100:200]):.2f}")

One sentence: 3,000 rows spun out of 30 points carry exactly the diversity of those 30 -- all huddled around cluster A -- so cluster B stays invisible to the model and recall there falls off a cliff, which is the whole point: volume is not variety. Right, homework done. Now let us talk about what any of this costs.

The GPU landscape: what things actually cost

Let us start with concrete numbers, because "GPUs are expensive" is not a plan. GPU pricing swings wildly on provider, instance type, commitment level and region. Here are approximate on-demand hourly rates for popular instances as of early 2026 (prices move constantly -- treat these as a map, not a quote, and verify before you commit real money):

NVIDIA A100 80GB (the reliable workhorse):

  • AWS p4d.24xlarge (8x A100): ~$32/hr
  • GCP a2-ultragpu-8g (8x A100): ~$29/hr
  • Azure ND96amsr_A100_v4 (8x A100): ~$27/hr
  • Lambda Cloud (8x A100): ~$10/hr

NVIDIA H100 80GB (the current king):

  • AWS p5.48xlarge (8x H100): ~$98/hr
  • GCP a3-highgpu-8g (8x H100): ~$88/hr
  • CoreWeave (8x H100): ~$24/hr

NVIDIA L4 / A10G (inference and light training):

  • AWS g6.xlarge (1x L4): ~$0.80/hr
  • GCP g2-standard-4 (1x L4): ~$0.70/hr

Consumer GPUs, for reference:

  • RTX 4090 24GB: ~$1,600 one-time
  • RTX 3090 24GB: ~$800 used

Look at that spread. Lambda charges roughly a third of what AWS charges for equivalent A100 compute. The smaller GPU-focused shops (Vast.ai, RunPod, Together AI) can be cheaper still, sometimes renting out community-sourced GPUs at bargain rates -- with matching tradeoffs in reliability and, frankly, security (somebody else's machine is exactly that). The lesson is blunt: the same silicon can cost you 3-5x more depending purely on whose logo is on the invoice. Comparison shopping is not being cheap, it is being competent.

Here is a tiny estimator so you can stop guessing and start computing:

def estimate_training_cost(gpu_hours, gpu_type="A100", provider="aws", use_spot=False):
    # Approximate per-GPU hourly rates (on-demand, early 2026)
    rates = {
        ("A100", "aws"): 4.00, ("A100", "gcp"): 3.67, ("A100", "lambda"): 1.25,
        ("H100", "aws"): 12.25, ("H100", "gcp"): 11.00, ("H100", "coreweave"): 3.00,
        ("L4", "aws"): 0.80, ("L4", "gcp"): 0.70,
    }
    rate = rates.get((gpu_type, provider), 4.00)
    if use_spot:
        rate *= 0.3                                  # spot ~ 60-70% cheaper
    total = gpu_hours * rate
    print(f"{gpu_hours:>5.0f} GPU-h  {provider:>9} {gpu_type:<5} "
          f"${total:>8,.0f}  ({'spot' if use_spot else 'on-demand'})")
    return total

# Fine-tuning a 7B model: ~50 GPU-hours on a single A100
estimate_training_cost(50, "A100", "aws",    use_spot=False)   # ~$200
estimate_training_cost(50, "A100", "aws",    use_spot=True)    # ~$60
estimate_training_cost(50, "A100", "lambda", use_spot=False)   # ~$63

# Training a custom vision model: ~500 GPU-hours
estimate_training_cost(500, "A100", "aws",    use_spot=True)    # ~$600
estimate_training_cost(500, "A100", "lambda", use_spot=False)   # ~$625

Notice something: AWS-on-spot and Lambda-on-demand land in roughly the same place. There is rarely one right answer -- there are several nearly-equivalent ones and a handful of catastrophically wrong ones. Your job is to avoid the second category.

Spot instances: the single biggest lever

Spot instances (AWS), preemptible VMs (GCP) and low-priority VMs (Azure) hand you the exact same hardware at a 60-90% discount. The catch, and there is always a catch: the provider can yank the machine back with little warning (usually a 2-minute notice) when an on-demand customer wants it.

That sounds terrifying until you realize training is one of the most interruptible workloads that exists -- if your code can checkpoint. The trick is to save state often enough that losing a machine costs you minutes, not days:

import torch, os, signal, sys

class CheckpointManager:
    """Save state regularly so a spot interruption never loses real progress."""
    def __init__(self, save_dir, save_every_n_steps=500):
        self.save_dir, self.save_every, self.step = save_dir, save_every_n_steps, 0
        os.makedirs(save_dir, exist_ok=True)
        signal.signal(signal.SIGTERM, self._emergency_save)   # AWS sends SIGTERM on reclaim

    def _emergency_save(self, signum, frame):
        print("SPOT TERMINATION NOTICE -- flushing checkpoint...")
        self.save(self.model, self.optimizer, emergency=True)
        sys.exit(0)

    def maybe_save(self, model, optimizer, loss):
        self.model, self.optimizer = model, optimizer          # keep refs for the emergency save
        self.step += 1
        if self.step % self.save_every == 0:
            self.save(model, optimizer)

    def save(self, model, optimizer, emergency=False):
        tag = "emergency" if emergency else f"step_{self.step}"
        path = os.path.join(self.save_dir, f"checkpoint_{tag}.pt")
        torch.save({"step": self.step,
                    "model_state_dict": model.state_dict(),
                    "optimizer_state_dict": optimizer.state_dict()}, path)
        print(f"saved {path}")

    def load_latest(self, model, optimizer):
        cks = sorted((f for f in os.listdir(self.save_dir) if f.endswith(".pt")),
                     key=lambda f: os.path.getmtime(os.path.join(self.save_dir, f)))
        if not cks:
            return 0
        state = torch.load(os.path.join(self.save_dir, cks[-1]), weights_only=False)
        model.load_state_dict(state["model_state_dict"])
        optimizer.load_state_dict(state["optimizer_state_dict"])
        print(f"resumed at step {state['step']}")
        return state["step"]

The pattern is simple: checkpoint to persistent storage (S3, GCS, a network drive -- not the ephemeral instance disk, which vanishes with the machine), and when the spot node dies, a fresh one resumes from the last checkpoint. You lose a couple of minutes and you save two-thirds of the bill. Having said that, the fear is mostly overblown in practice -- GPU spot instances on AWS often run for hours or days before anyone reclaims them. GCP preemptibles have a hard 24-hour ceiling, but Google also offers "spot VMs" that can run longer. Build the interruption handling once, test it once, and then genuinely never think about it again.

On-premise versus cloud: the crossover point

The whole cloud-versus-own-hardware debate collapses into one word: utilization. Cloud is pay-per-use. On-premise is pay-a-lot-upfront-then-nearly-free.

A single RTX 4090 costs ~$1,600. An equivalent A10G instance on AWS runs ~$1.00/hr on-demand. Naively, your 4090 "pays for itself" after about 1,600 hours -- roughly 67 days of 24/7 grind, or eight months at eight hours a day. But that naive number ignores electricity (a 4090 pulls ~450W under load, at maybe $0.10-0.30/kWh), cooling, the desk it sits on, and the hours of your life spent being its sysadmin. Fold those in and breakeven stretches to something like 12-18 months.

def on_prem_vs_cloud(gpu_cost, electricity_per_kwh, gpu_watts,
                     cloud_hourly, hours_per_week, years=3):
    """Total cost of ownership over N years -- the honest version."""
    total_hours = hours_per_week * 52 * years
    electricity = (gpu_watts / 1000) * electricity_per_kwh * total_hours
    on_prem = gpu_cost + electricity
    cloud = cloud_hourly * total_hours
    print(f"--- {years}yr @ {hours_per_week}h/week ---")
    print(f"on-premise: ${on_prem:>8,.0f}  (hw ${gpu_cost:,.0f} + power ${electricity:,.0f})")
    print(f"cloud:      ${cloud:>8,.0f}")
    print(f"winner:     {'on-premise' if on_prem < cloud else 'cloud'}\n")
    return on_prem, cloud

on_prem_vs_cloud(1600, 0.15, 450, 1.00, 40, years=3)   # heavy use: 40h/week
on_prem_vs_cloud(1600, 0.15, 450, 1.00, 10, years=3)   # light use: 10h/week

Run it and the answer flips on you: at 40 hours a week the 4090 wins in a landslide, at 10 hours a week the cloud wins because you are not paying to keep idle silicon warm. This is why serious teams almost always land on a hybrid shape -- owned GPUs for the steady, predictable load (daily inference, the nightly retrain), cloud burst capacity for the spiky experiments and the occasional monster training run. You minimize both idle hardware and idle wallet.

Cost-optimization strategies that do not cost you quality

Beyond picking providers and instances, a handful of techniques quietly shrink the bill without touching your final metrics:

Start small, scale up. Debug your pipeline on a 10% sample with a smaller architecture. Verify the data flows, the loss goes down, the hyperparameters are in the right ballpark. Then go full scale, once. One full-scale run you know will succeed is far cheaper than ten exploratory full-scale runs you launched on hope.

Mixed precision. We met this back in episode #43 -- float16 or bfloat16 for the bulk of the arithmetic roughly halves memory, which lets you rent a smaller (cheaper) GPU or pack a bigger batch onto the one you have. On modern tensor-core hardware it is also faster, so you are billed for fewer hours. Free money, basically.

Gradient accumulation. Cannot afford eight GPUs for a big effective batch? Accumulate gradients over several small batches (episode #126) and simulate the big one on cheaper hardware. You trade wall-clock time for dollars -- slower, but cheaper per run.

Distillation. Train the big expensive model once, then distill it into a small cheap one for inference (episode #120). The teacher might cost $10,000 to train, but if the student shaves $5,000/month off your serving bill, it has paid for itself before the quarter is out.

Cache the boring parts. Do not recompute embeddings or frozen-backbone features every single run. Precompute once, dump to disk, load from disk. When you are fine-tuning on top of a frozen backbone, precomputing its outputs can cut training time by 80% -- you are literally paying to run the same forward pass over and over otherwise.

The make-versus-buy decision

This one shows up at every level. Train your own model or call an API? Roll your own infra or rent a managed service? Write a training framework or use one that exists? The instinct of the proud engineer is always "build it myself" -- and that instinct is often expensive nonsense ;-)

Call an API (OpenAI, Anthropic, Google) when your task is well-served by a general model, your latency budget is relaxed (100ms+ is fine), your data is allowed to leave the building, and shipping speed matters more than per-query cost. Train your own when you need domain expertise the general models lack, when your query volume is high enough that API fees dwarf training costs, when privacy law forces the data to stay on your own iron, or when you need real control over how and when the model changes.

The crossover is just arithmetic, so do the arithmetic:

def api_vs_custom(queries_per_month, api_cost_per_query,
                  training_cost, inference_cost_per_month, months=12):
    """When does owning the model beat renting the API?"""
    api_total = queries_per_month * api_cost_per_query * months
    custom_total = training_cost + inference_cost_per_month * months
    breakeven = next((m for m in range(1, months + 1)
                      if training_cost + inference_cost_per_month * m
                      < queries_per_month * api_cost_per_query * m), None)
    print(f"--- {months}-month horizon ---")
    print(f"API total:    ${api_total:>9,.0f}")
    print(f"custom total: ${custom_total:>9,.0f}")
    print(f"custom breaks even at month {breakeven}" if breakeven
          else "API stays cheaper over this horizon")

# 100k queries/month at $0.03 each  vs.  $5k training + $500/mo hosting
api_vs_custom(100_000, 0.03, 5000, 500, months=12)

Play with the inputs. At low volume the API wins forever; crank the query count up and there is always a month where owning the thing overtakes renting it. The point is not the specific number -- it is that you can know the number instead of arguing about it in a meeting.

Budgeting an AI project across its whole life

If you are scoping a project -- for yourself, a client, or a pitch deck -- do not budget it as one lump labelled "training". Budget it in phases, because AI projects do not end when the loss curve flattens. They end when the thing runs reliably, at acceptable cost, in production.

def budget_plan(total_budget):
    phases = {"exploration": 0.10, "development": 0.25,
              "final training": 0.20, "inference (ongoing)": 0.45}
    print(f"total budget: ${total_budget:,.0f}")
    for name, frac in phases.items():
        print(f"  {name:<22} ${total_budget * frac:>9,.0f}  ({frac:>4.0%})")

budget_plan(50_000)

Exploration (5-15%). Small experiments, data spelunking, throwaway prototypes. Smallest GPU that fits, or plain CPU for the early data work. This phase answers one question: is the problem even solvable?

Development (20-35%). The real training, hyperparameter search, architecture bake-offs. This is where spot instances earn their keep -- run many experiments, but keep them small.

Final training (15-25%). Train your best config at full scale, once (or a small handful of times). On-demand is defensible here because you are paying for reliability, not exploration.

Inference (30-50%, forever). Serving the model in production -- usually the biggest line item over the project's life, and the one people forget entirely. Optimize it hard: quantization, batching, caching, right-sized instances. The classic, fatal mistake is blowing 80% of the budget on training experiments and having nothing left to actually run the model for users.

Hidden costs that quietly blow the budget

A few line items that never make it into the plan but always make it into the invoice:

Data transfer. Shuffling data between storage and instances, across regions, or out of the cloud entirely. AWS charges ~$0.09/GB egress -- move a 500GB dataset around carelessly and that is $45 a pop, every pop. Serving model outputs to users at scale adds up faster than anyone expects.

Storage. Checkpoints, logs, datasets, intermediate junk. One large training run can spit out hundreds of gigabytes of checkpoints. At ~$0.023/GB/month on S3 it is not devastating, but it accumulates, and nobody ever goes back to delete last quarter's experiments.

Idle instances. The p4d you forgot to shut down on Friday that hummed away all weekend doing precisely nothing. At ~$32/hr that is $1,536 for a weekend of zero work -- the single most preventable overrun in all of cloud AI. So prevent it:

def auto_shutdown_guard(utilization_history, low_threshold=0.05, patience=6):
    """Return True (=> shut the box down) after `patience` consecutive idle checks."""
    idle_streak = 0
    for util in utilization_history:                 # each entry: recent GPU utilization 0..1
        idle_streak = idle_streak + 1 if util < low_threshold else 0
        if idle_streak >= patience:
            return True
    return False

# six straight near-zero readings -> the guard pulls the plug before the weekend bill lands
readings = [0.72, 0.68, 0.01, 0.00, 0.02, 0.00, 0.01, 0.00]
print("SHUTDOWN" if auto_shutdown_guard(readings) else "keep running")

Wire that into a cron job with real utilization numbers (nvidia-smi will hand them to you) and a shutdown call, set a billing alert as a second line of defense, and that particular category of pain simply stops happening.

Engineering time. The cheapest GPU-hour is the one you never needed because your code was efficient. Spend two days making a training pipeline run twice as fast and you have halved the compute bill for every run that follows, forever. Engineer time is expensive -- but GPU time, compounded over a project, is frequently more expensive. That is not an excuse to gold-plate everything (I might suffer from a touch of 'perfectionism' myself, so I know the temptation), it is a reminder that optimization has a real, calculable payoff.

What you should remember

  • GPU pricing varies 3-5x across providers for the same hardware -- always compare before committing a cent;
  • spot / preemptible instances save 60-90% and are practical for any training job the moment your code checkpoints properly;
  • on-premise beats cloud once utilization climbs past roughly 30-40% of available hours over a couple of years -- and hybrid usually beats both;
  • cost optimization starts before training: start small, validate the approach, then scale the thing you already know works;
  • make-versus-buy is arithmetic, not ideology -- compute the crossover from query volume, privacy needs, and how well a general model serves you;
  • budget across four phases -- exploration, development, final training, inference -- and remember inference is usually the biggest long-term cost;
  • the hidden costs (egress, idle boxes, storage) silently double bills, and almost every one of them is preventable with a billing alert and a shutdown script.

Exercises

Three to gnaw on before next time -- and, as always, I will walk through full solutions at the top of the following episode.

  1. Turn the estimator into a budget tool. Extend estimate_training_cost so it also takes a hard budget ceiling and, for each provider/instance in the rates table, prints how many GPU-hours that budget buys you (on-demand and spot). Then have it declare which single option lets you train for the longest under the ceiling. One sentence on whether the "longest" option and the "cheapest per hour" option are ever different -- and why they can be.

  2. Model your own crossover. Pick a realistic monthly query volume for a project you might actually build, plug in today's published API price for a model of your choice (go look it up), and use api_vs_custom to find the breakeven month. Now redo it assuming your traffic grows 15% month-over-month instead of staying flat, and report how much earlier (or later) owning the model starts to pay off.

  3. Make the idle guard real-ish. Rewrite auto_shutdown_guard so that instead of a fixed list it consumes a stream of simulated readings (draw utilization from numpy's random generator, with an occasional burst of real work), runs for a simulated "week" of checks every few minutes, and reports the total dollars it saved versus a box left running the whole week at $32/hr. One sentence on how you would choose patience so you do not kill a job that is merely between epochs.

We have spent this whole episode reasoning about dollars, watts and GPU-hours -- treating infrastructure as a thing you can calculate your way through. And you can. But step back and notice what every one of these decisions quietly assumed: that someone decides which provider, someone owns the shutdown script, someone chooses build-versus-buy and lives with it. Most AI projects, in my experience, do not die on the math or the GPUs -- they die in the gaps between people, on the handoffs nobody owned and the decisions nobody was actually allowed to make. Compute you can budget on a spreadsheet. The harder question I want you carrying into next time is the one the spreadsheet cannot answer: how do you organize the humans around all this so the money you just learned to save does not leak straight back out through process? That is where we head next ;-)

Bedankt en tot de volgende keer -- now go read your cloud bill before it reads you!

scipio@scipio

Learn AI Series (#134) - AI Infrastructure Economics | Ecency