# Checkpointing a run

Checkpoints are what Artifacts was built for: written often, read rarely, almost
always by the next box of the same run. For a training run the only change is
the location string: `s3://bucket/runs/my-run` becomes
`o41://experiments/my-run`.

## SexpGPU

> SexpGPU takes `o41://` from its release with o41-checkpoint, which is in
> review. Until then, use the [o41-checkpoint](https://artifacts.041.io/docs/o41-checkpoint.md) crate, the
> [CLI](https://artifacts.041.io/docs/cli.md) or the [HTTP API](https://artifacts.041.io/docs/api.md).

```bash
export O41_ARTIFACTS_API_KEY=ak_…
SEXPGPU_CHECKPOINT_DIR=o41://experiments/run-1234 \
  sexpgpu run train.sx --resume latest
```

- `--resume latest` opens the artifact from the latest, or from nothing on the
  first launch. Put it on the command line from the first launch, so every
  relaunch is the same command.
- No `--resume` starts fresh and is refused on an artifact that has checkpoints.
- `--resume o41://experiments/run-1234/step-00000300` forks: the checkpoint
  location must be a new, empty artifact.
- The key is checked before step 1, so a run never finds out hours in that it
  cannot write.

## A spot run hopping clouds

1. Box A on **aws us-east-1** writes `step-00000100` … `step-00000500` into your
   S3 bucket in us-east-1. It is preempted.
2. SkyPilot brings the job back on **gcp europe-west4**. Same command. Before
   the first step the run opens the artifact from the latest, with its place
   taken from `SKYPILOT_CLUSTER_INFO`.
3. The open fences box A, should it still be alive, and returns step 500 signed,
   marked `cross-cloud`. The run reads it with parallel ranged GETs: the one
   cross-cloud transfer.
4. `step-00000510` and on land in your GCS bucket in europe-west4: a region
   match, free.
5. Preempted again, it comes back on **azure westeurope**, reads step 900 once,
   and writes to your Blob container there.

A box in a cloud you have no location in still trains: its checkpoints go to the
cloud's default location, or to your fallback, with a warning the run logs. The
usage page shows where that happens, which is where a new location pays.

## Any other trainer

The [o41-checkpoint](https://artifacts.041.io/docs/o41-checkpoint.md) crate is the same thing for any Rust
program; the [HTTP API](https://artifacts.041.io/docs/api.md) is the contract for every other language. A
Python eval script needs one `versions/read` and plain HTTP GETs:

```python
import os, requests

api = "https://artifacts.041.io/api/v1"
auth = {"Authorization": f"Bearer {os.environ['O41_ARTIFACTS_API_KEY']}"}
v = requests.get(f"{api}/versions/read", headers=auth, params={
    "artifact": "experiments/run-1234", "version": "latest", "urls": "1",
}).json()
for f in v["files"]:
    with open(f["name"], "wb") as out:
        out.write(requests.get(f["url"]).content)
```

## What it costs

Writes are the volume, and they are free when local. A 20 GB checkpoint every 30
minutes for a week is about 6.7 TB written: $0 into the box's own region, about
$130 to another region of the same cloud, $550 to $800 out of the cloud. One
cross-cloud resume of the latest is about $2. Nothing is replicated ahead of
time.

---

Artifacts by 041 documentation. Every page: https://artifacts.041.io/llms.txt
