Get started
Managed GPU jobs on Vast.ai

Cheap GPUs that don't lose your progress.

Community GPUs cost a fraction of AWS, but machines disappear mid-run. NomadGPU picks reliable hosts, saves your checkpoints every few minutes and moves the job to a new machine when one fails — automatically, with hard spending limits.

QLoRA Qwen2.5-7B, 300 steps: $0.19 on one RTX 4090 vs $1.35–1.80 on an AWS A100. Method →

Example of what you get in Telegram when a machine fails. Times are illustrative.

How it works

You send the job. We keep it alive.

Vast.ai rents GPUs from independent hosts at a fraction of big-cloud prices. The catch: hosts go offline, spot machines get outbid, slow disks and bad networks waste hours. That is the part we handle.

1

Pick a good host

Not the cheapest one: reliability, network, PCIe and disk speed, and the traffic price — which can cost more than the GPU itself.

2

Save progress

An agent inside the container uploads your checkpoints to storage every few minutes and reports a heartbeat every minute.

3

Move on failure

Host lost or frozen? We rent a replacement and resume from the last checkpoint — usually within minutes, no action from you.

4

Stop at the limit

Per-job and monthly budgets. Warning at 80%, at 100% the machine is released and progress is kept.

Benchmark

One real job, billed numbers.

QLoRA fine-tune of Qwen2.5-7B-Instruct, 300 steps. Cost from the Vast.ai invoice, not from "price × time".

86–89%

cheaper on this job than one A100 in AWS

RTX 4090 on Vast.ai, managed by NomadGPU$0.19
AWS A100, if it were ~1.5× faster$1.35
AWS A100, same time$1.80
Host1× RTX 4090, reliability 0.996, $0.43/h incl. traffic
Timeagent online in 58 s, training 18.8 min, 0.44 h in total
Result300 steps, final loss 1.04, 5 of 5 checkpoint syncs OK
Vast.ai bill≈ $0.19 (GPU, disk and traffic)
AWS baselinep4d.24xlarge on-demand per A100: $4.10/h
Method and caveats

Public image pytorch/pytorch:2.5.1-cuda12.4, dependencies installed at start, model download (~15 GB) included in the bill. Dataset yahma/alpaca-cleaned, 4-bit QLoRA, checkpoint every 50 steps.

The A100 is faster on this task; we estimate 1.3–1.5× (estimate, not measured), hence the $1.35–1.80 range. AWS does not rent a single A100 on demand — p4d is 8 GPUs, so the per-GPU price favours AWS.

An earlier run on a host with expensive traffic cost $0.81 — traffic ate most of the savings. That is why host selection now ranks by the full hourly price including traffic. Other jobs will save a different amount; this is one measured job, not a promise.

Demo

We kill the machine mid-training. The job keeps going.

Two minutes, no voice-over: start a fine-tune, destroy the host by hand, watch the job resume from its checkpoint on another machine.

Video coming soon.
Want to see it live? Ask for a demo run — it costs about $0.05.
Your account, your money

We never hold your GPU budget.

Jobs run on your own Vast.ai account. You pay Vast directly; our fee covers only the service. Revoke our key at any moment.

Your Vast account

You top up your own balance. We get an API key named nomadgpu — delete it and we are out.

Hard limits

Monthly and per-job budgets. At 100% jobs stop and keep their progress. Raising a limit is one message.

Isolated storage

Each job gets temporary storage keys limited to its own folder for 4 hours. Checkpoints can be encrypted with your key.

What we take care of

  • Choosing reliable, fairly priced machines
  • Starting jobs, checkpoints, moving on failure
  • Spending limits and spending reports
  • Telegram alerts and a private status page

What stays on your side

  • Your code: if the script has a bug, we show you the log and the error
  • Confidential data: hosts are independent, don't run personal or medical data this way
  • GPU bills: paid by you, directly to Vast.ai
Pricing

Simple service fee. GPUs at cost, on your account.

No prepayment: you are billed after the period. Cancel anytime.

For production workloads

Managed

from $300 / month
for GPU bills from ~$1,500 / month · setup quoted per pipeline
  • Your own pipelines and images, built with you
  • Inference services kept running 24/7
  • Fast reaction, including nights
  • Monthly savings report against your old cloud
Talk to us
FAQ

Questions

What do I need to start?

A Vast.ai account with some credit (a 7B LoRA fine-tune costs roughly $0.20–0.50 in GPU time), an API key for us, and your job: a Docker image and a command, or one of our templates. Setup takes about 10 minutes on your side.

What happens when a machine fails?

We rent a replacement automatically and the job continues from its last checkpoint, usually within a few minutes (the model may need to download again on the new machine). You get a message when the machine is lost and when the job is running again.

Does my code need changes?

Only one: save checkpoints to the folder we give you (an environment variable) and resume from it when it is not empty. Hugging Face Trainer does this with a single argument. We help with this during setup.

Can the machine owner see my data?

In principle, yes: Vast machines belong to independent hosts and your job runs there unencrypted. Checkpoints in storage can be encrypted, but do not process personal, medical or confidential data this way.

Is it slower than AWS?

Training speed per GPU-hour on an RTX 4090 is close to an A100 for small and medium models. Start-up is slower: a new machine needs a few minutes to download the image and the model.

Which GPU providers do you support?

Vast.ai today. RunPod, Lambda Cloud and TensorDock are planned; the model is the same — your account, our management.

How do I stop?

Tell us, and delete the nomadgpu key in your Vast account. Your checkpoints stay in storage for you.

Get started

Send us your job. We'll run it on a cheap GPU this week.

Tell us the model, the GPU you use now and roughly what you spend per month. We reply within one business day with a host estimate and a plan.