No Docker. No Kubernetes.
No cloud setup. No YAML.
Bring any model from HuggingFace or your own weights. Fleet picks the cheapest GPU that fits and bills by the second.
Every team rebuilds the same plumbing to get a model onto a GPU and behind an API. Fleet is the abstraction that makes all of it disappear.
Weights in, endpoint out. Fleet hides the GPUs, containers, and schedulers behind plain Python objects.
No Docker, no Kubernetes, no cloud console. Nothing to provision, patch, or page you at night.
Stay in your editor. The same script that runs on your laptop runs on a fleet of cloud GPUs.
Trained model to live, streaming endpoint in one call, and you only pay for the seconds you use.
Mission control for your models. Kick off training runs, spin up deployments, and test live inference, all from one neon cockpit, no terminal required.
Loss, learning rate, validation, streamed live across thousands of steps. Toggle a trendline to see through the noise and catch a run going sideways before it costs you GPU-hours.
Dashboard, SDK, CLI, or autonomous agent: Fleet is the same platform from every angle, built to be driven by people and machines alike.
Every page is served as clean Markdown an LLM can ingest. No scraping, no HTML soup.
One-click "copy for LLM" on every doc, plus an MCP endpoint so coding agents can call Fleet directly.
The CLI, the SDK, and the docs all speak the same shapes. What a human reads is what an agent runs.
No seats. No subscriptions. No idle servers draining your balance. Pick the tier your model needs and pay by the second it runs.
Inference runs on serverless GPU infrastructure. It scales up with demand and back down to zero when idle, so you only pay while requests are being served.