NO OPS!

TRAIN & DEPLOYON CLOUD GPUSFROM LOCAL PYTHON.

No Docker. No Kubernetes.
No cloud setup. No YAML.

Bring any model from HuggingFace or your own weights. Fleet picks the cheapest GPU that fits and bills by the second.

GET STARTED →
READ THE DOCS
// LAUNCH SEQUENCE
FLEET // PYTHON
// THE MISSING PIECE

The gap between a model on your laptop and a model in production is all ops. Fleet is that piece.

Every team rebuilds the same plumbing to get a model onto a GPU and behind an API. Fleet is the abstraction that makes all of it disappear.

WITHOUT FLEET
  • Write a Dockerfile
  • Provision GPU instances
  • Configure Kubernetes
  • Wire up cloud IAM & VPCs
  • Install CUDA & drivers
  • Build a job scheduler
  • Stand up a serving layer
  • Babysit it all at 3am
WITH FLEET
model = fleet.load("./mistral")
# upload weights
job = model.train("./data")
# fine-tune on GPU
endpoint = job.model().deploy()
# live in seconds
That's the whole stack.
ONE ABSTRACTION

Weights in, endpoint out. Fleet hides the GPUs, containers, and schedulers behind plain Python objects.

ZERO INFRASTRUCTURE

No Docker, no Kubernetes, no cloud console. Nothing to provision, patch, or page you at night.

LOCAL-FIRST WORKFLOW

Stay in your editor. The same script that runs on your laptop runs on a fleet of cloud GPUs.

SECONDS, NOT SPRINTS

Trained model to live, streaming endpoint in one call, and you only pay for the seconds you use.

// PREFER A DASHBOARD?
Fleet dashboard showing a model training on an A100 GPU

Not just a stack of boring tables

Mission control for your models. Kick off training runs, spin up deployments, and test live inference, all from one neon cockpit, no terminal required.

Train modelsSpin up deploysTest inferenceWatch the GPUs
OPEN THE DASHBOARD →
// LIVE METRICS

Monitor the stuff that matters

Loss, learning rate, validation, streamed live across thousands of steps. Toggle a trendline to see through the noise and catch a run going sideways before it costs you GPU-hours.

LossLearning rateVal lossTrendlines
Training metrics over a 10,000-step run: loss curve with learning-rate trendline
// FOR HUMANS & MACHINES

Live in the terminal? There's a CLI. Building agents? The docs read themselves.

Dashboard, SDK, CLI, or autonomous agent: Fleet is the same platform from every angle, built to be driven by people and machines alike.

THE CLI
$ fleet train ./model --data ./data.jsonl
→ job_b1b9… scheduled on H100 · streaming logs
$ fleet deploy job_b1b9 --gpu a100
→ endpoint live · https://api.fleet.dev/…
$ fleet infer endpoint_7c2 "Hello!"
AGENT-READY DOCS
llms.txt + Markdown docs

Every page is served as clean Markdown an LLM can ingest. No scraping, no HTML soup.

Copy-paste for agents

One-click "copy for LLM" on every doc, plus an MCP endpoint so coding agents can call Fleet directly.

Same API, both worlds

The CLI, the SDK, and the docs all speak the same shapes. What a human reads is what an agent runs.

// PRICING

Pay for what you actually use.

No seats. No subscriptions. No idle servers draining your balance. Pick the tier your model needs and pay by the second it runs.

How many requests?
1K10K100K1M10M
Standard · 100K requests
$45.90≈ ~1.2s/req est.
Estimated from typical request latency. Actual cost depends on your model and response length. Scales to zero, so idle endpoints cost nothing.
START FREE →

Inference runs on serverless GPU infrastructure. It scales up with demand and back down to zero when idle, so you only pay while requests are being served.

READY TO LAUNCH?

Join the engineers already training on Fleet.

GET STARTED FOR FREE →
★ FLEET PLATFORM ONLINE★ GPU SCHEDULING ACTIVE★ ALL SYSTEMS GO★ INFERENCE ENDPOINTS READY★ MISSION CONTROL STANDING BY★ LAUNCH IN T-MINUS ZERO★ FLEET PLATFORM ONLINE★ GPU SCHEDULING ACTIVE★ ALL SYSTEMS GO★ INFERENCE ENDPOINTS READY★ MISSION CONTROL STANDING BY★ LAUNCH IN T-MINUS ZERO★ FLEET PLATFORM ONLINE★ GPU SCHEDULING ACTIVE★ ALL SYSTEMS GO★ INFERENCE ENDPOINTS READY★ MISSION CONTROL STANDING BY★ LAUNCH IN T-MINUS ZERO★ FLEET PLATFORM ONLINE★ GPU SCHEDULING ACTIVE★ ALL SYSTEMS GO★ INFERENCE ENDPOINTS READY★ MISSION CONTROL STANDING BY★ LAUNCH IN T-MINUS ZERO★ FLEET PLATFORM ONLINE★ GPU SCHEDULING ACTIVE★ ALL SYSTEMS GO★ INFERENCE ENDPOINTS READY★ MISSION CONTROL STANDING BY★ LAUNCH IN T-MINUS ZERO★ FLEET PLATFORM ONLINE★ GPU SCHEDULING ACTIVE★ ALL SYSTEMS GO★ INFERENCE ENDPOINTS READY★ MISSION CONTROL STANDING BY★ LAUNCH IN T-MINUS ZERO