Cheap GPUs that don't lose your progress.
Community GPUs cost a fraction of AWS, but machines disappear mid-run. NomadGPU picks reliable hosts, saves your checkpoints every few minutes and moves the job to a new machine when one fails — automatically, with hard spending limits.
QLoRA Qwen2.5-7B, 300 steps: $0.19 on one RTX 4090 vs $1.35–1.80 on an AWS A100. Method →
Example of what you get in Telegram when a machine fails. Times are illustrative.
You send the job. We keep it alive.
Vast.ai rents GPUs from independent hosts at a fraction of big-cloud prices. The catch: hosts go offline, spot machines get outbid, slow disks and bad networks waste hours. That is the part we handle.
Pick a good host
Not the cheapest one: reliability, network, PCIe and disk speed, and the traffic price — which can cost more than the GPU itself.
Save progress
An agent inside the container uploads your checkpoints to storage every few minutes and reports a heartbeat every minute.
Move on failure
Host lost or frozen? We rent a replacement and resume from the last checkpoint — usually within minutes, no action from you.
Stop at the limit
Per-job and monthly budgets. Warning at 80%, at 100% the machine is released and progress is kept.
One real job, billed numbers.
QLoRA fine-tune of Qwen2.5-7B-Instruct, 300 steps. Cost from the Vast.ai invoice, not from "price × time".
cheaper on this job than one A100 in AWS
| Host | 1× RTX 4090, reliability 0.996, $0.43/h incl. traffic |
|---|---|
| Time | agent online in 58 s, training 18.8 min, 0.44 h in total |
| Result | 300 steps, final loss 1.04, 5 of 5 checkpoint syncs OK |
| Vast.ai bill | ≈ $0.19 (GPU, disk and traffic) |
| AWS baseline | p4d.24xlarge on-demand per A100: $4.10/h |
Method and caveats
Public image pytorch/pytorch:2.5.1-cuda12.4, dependencies installed at start, model download (~15 GB) included in the bill. Dataset yahma/alpaca-cleaned, 4-bit QLoRA, checkpoint every 50 steps.
The A100 is faster on this task; we estimate 1.3–1.5× (estimate, not measured), hence the $1.35–1.80 range. AWS does not rent a single A100 on demand — p4d is 8 GPUs, so the per-GPU price favours AWS.
An earlier run on a host with expensive traffic cost $0.81 — traffic ate most of the savings. That is why host selection now ranks by the full hourly price including traffic. Other jobs will save a different amount; this is one measured job, not a promise.
We kill the machine mid-training. The job keeps going.
Two minutes, no voice-over: start a fine-tune, destroy the host by hand, watch the job resume from its checkpoint on another machine.
Want to see it live? Ask for a demo run — it costs about $0.05.
We never hold your GPU budget.
Jobs run on your own Vast.ai account. You pay Vast directly; our fee covers only the service. Revoke our key at any moment.
Your Vast account
You top up your own balance. We get an API key named nomadgpu — delete it and we are out.
Hard limits
Monthly and per-job budgets. At 100% jobs stop and keep their progress. Raising a limit is one message.
Isolated storage
Each job gets temporary storage keys limited to its own folder for 4 hours. Checkpoints can be encrypted with your key.
What we take care of
- Choosing reliable, fairly priced machines
- Starting jobs, checkpoints, moving on failure
- Spending limits and spending reports
- Telegram alerts and a private status page
What stays on your side
- Your code: if the script has a bug, we show you the log and the error
- Confidential data: hosts are independent, don't run personal or medical data this way
- GPU bills: paid by you, directly to Vast.ai
Simple service fee. GPUs at cost, on your account.
No prepayment: you are billed after the period. Cancel anytime.
Starter
- Setup from a ready template (LoRA, vLLM, ComfyUI)
- Automatic recovery from checkpoints
- Budgets, Telegram alerts, status page
- Email support, next business day
Managed
- Your own pipelines and images, built with you
- Inference services kept running 24/7
- Fast reaction, including nights
- Monthly savings report against your old cloud
Questions
What do I need to start?
A Vast.ai account with some credit (a 7B LoRA fine-tune costs roughly $0.20–0.50 in GPU time), an API key for us, and your job: a Docker image and a command, or one of our templates. Setup takes about 10 minutes on your side.
What happens when a machine fails?
We rent a replacement automatically and the job continues from its last checkpoint, usually within a few minutes (the model may need to download again on the new machine). You get a message when the machine is lost and when the job is running again.
Does my code need changes?
Only one: save checkpoints to the folder we give you (an environment variable) and resume from it when it is not empty. Hugging Face Trainer does this with a single argument. We help with this during setup.
Can the machine owner see my data?
In principle, yes: Vast machines belong to independent hosts and your job runs there unencrypted. Checkpoints in storage can be encrypted, but do not process personal, medical or confidential data this way.
Is it slower than AWS?
Training speed per GPU-hour on an RTX 4090 is close to an A100 for small and medium models. Start-up is slower: a new machine needs a few minutes to download the image and the model.
Which GPU providers do you support?
Vast.ai today. RunPod, Lambda Cloud and TensorDock are planned; the model is the same — your account, our management.
How do I stop?
Tell us, and delete the nomadgpu key in your Vast account. Your checkpoints stay in storage for you.
Send us your job. We'll run it on a cheap GPU this week.
Tell us the model, the GPU you use now and roughly what you spend per month. We reply within one business day with a host estimate and a plan.