All pages
Checkpointing a run
Checkpoints are what Artifacts was built for: written often, read rarely, almost
always by the next box of the same run. For a training run the only change is
the location string: s3://bucket/runs/my-run becomes
o41://experiments/my-run.
SexpGPU
SexpGPU takes
o41://from its release with o41-checkpoint, which is in review. Until then, use the o41-checkpoint crate, the CLI or the HTTP API.
export O41_ARTIFACTS_API_KEY=ak_…
SEXPGPU_CHECKPOINT_DIR=o41://experiments/run-1234 \
sexpgpu run train.sx --resume latest--resume latestopens the artifact from the latest, or from nothing on the first launch. Put it on the command line from the first launch, so every relaunch is the same command.- No
--resumestarts fresh and is refused on an artifact that has checkpoints. --resume o41://experiments/run-1234/step-00000300forks: the checkpoint location must be a new, empty artifact.- The key is checked before step 1, so a run never finds out hours in that it cannot write.
A spot run hopping clouds
- Box A on aws us-east-1 writes
step-00000100…step-00000500into your S3 bucket in us-east-1. It is preempted. - SkyPilot brings the job back on gcp europe-west4. Same command. Before
the first step the run opens the artifact from the latest, with its place
taken from
SKYPILOT_CLUSTER_INFO. - The open fences box A, should it still be alive, and returns step 500 signed,
marked
cross-cloud. The run reads it with parallel ranged GETs: the one cross-cloud transfer. step-00000510and on land in your GCS bucket in europe-west4: a region match, free.- Preempted again, it comes back on azure westeurope, reads step 900 once, and writes to your Blob container there.
A box in a cloud you have no location in still trains: its checkpoints go to the cloud's default location, or to your fallback, with a warning the run logs. The usage page shows where that happens, which is where a new location pays.
Any other trainer
The o41-checkpoint crate is the same thing for any Rust
program; the HTTP API is the contract for every other language. A
Python eval script needs one versions/read and plain HTTP GETs:
import os, requests
api = "https://artifacts.041.io/api/v1"
auth = {"Authorization": f"Bearer {os.environ['O41_ARTIFACTS_API_KEY']}"}
v = requests.get(f"{api}/versions/read", headers=auth, params={
"artifact": "experiments/run-1234", "version": "latest", "urls": "1",
}).json()
for f in v["files"]:
with open(f["name"], "wb") as out:
out.write(requests.get(f["url"]).content)What it costs
Writes are the volume, and they are free when local. A 20 GB checkpoint every 30 minutes for a week is about 6.7 TB written: $0 into the box's own region, about $130 to another region of the same cloud, $550 to $800 out of the cloud. One cross-cloud resume of the latest is about $2. Nothing is replicated ahead of time.