Artifacts
All pages
Docs · StartMarkdown

Checkpointing a run

Checkpoints are what Artifacts was built for: written often, read rarely, almost always by the next box of the same run. For a training run the only change is the location string: s3://bucket/runs/my-run becomes o41://experiments/my-run.

SexpGPU

SexpGPU takes o41:// from its release with o41-checkpoint, which is in review. Until then, use the o41-checkpoint crate, the CLI or the HTTP API.

export O41_ARTIFACTS_API_KEY=ak_…
SEXPGPU_CHECKPOINT_DIR=o41://experiments/run-1234 \
  sexpgpu run train.sx --resume latest
  • --resume latest opens the artifact from the latest, or from nothing on the first launch. Put it on the command line from the first launch, so every relaunch is the same command.
  • No --resume starts fresh and is refused on an artifact that has checkpoints.
  • --resume o41://experiments/run-1234/step-00000300 forks: the checkpoint location must be a new, empty artifact.
  • The key is checked before step 1, so a run never finds out hours in that it cannot write.

A spot run hopping clouds

  1. Box A on aws us-east-1 writes step-00000100 … step-00000500 into your S3 bucket in us-east-1. It is preempted.
  2. SkyPilot brings the job back on gcp europe-west4. Same command. Before the first step the run opens the artifact from the latest, with its place taken from SKYPILOT_CLUSTER_INFO.
  3. The open fences box A, should it still be alive, and returns step 500 signed, marked cross-cloud. The run reads it with parallel ranged GETs: the one cross-cloud transfer.
  4. step-00000510 and on land in your GCS bucket in europe-west4: a region match, free.
  5. Preempted again, it comes back on azure westeurope, reads step 900 once, and writes to your Blob container there.

A box in a cloud you have no location in still trains: its checkpoints go to the cloud's default location, or to your fallback, with a warning the run logs. The usage page shows where that happens, which is where a new location pays.

Any other trainer

The o41-checkpoint crate is the same thing for any Rust program; the HTTP API is the contract for every other language. A Python eval script needs one versions/read and plain HTTP GETs:

import os, requests

api = "https://artifacts.041.io/api/v1"
auth = {"Authorization": f"Bearer {os.environ['O41_ARTIFACTS_API_KEY']}"}
v = requests.get(f"{api}/versions/read", headers=auth, params={
    "artifact": "experiments/run-1234", "version": "latest", "urls": "1",
}).json()
for f in v["files"]:
    with open(f["name"], "wb") as out:
        out.write(requests.get(f["url"]).content)

What it costs

Writes are the volume, and they are free when local. A 20 GB checkpoint every 30 minutes for a week is about 6.7 TB written: $0 into the box's own region, about $130 to another region of the same cloud, $550 to $800 out of the cloud. One cross-cloud resume of the latest is about $2. Nothing is replicated ahead of time.