Where should I rent GPUs for AI inference or fine-tuning?
For bursty inference, serverless GPU with per-second billing wins because you pay nothing while idle. For sustained training, reserved instances are far cheaper. Cold start time is the specification to compare first.
Surveyed 17 September 2026
Choose compute and GPUs
Open in Tools →- 01Lambda
GPU cloud offering on-demand instances, multi-GPU clusters, and reserved capacity for training and inference.
Developer Tools - 02RunPod
GPU cloud renting pods by the second and running autoscaling serverless GPU endpoints for training and batch jobs.
Developer Tools - 03Vast.ai
Marketplace where independent hosts rent out GPUs, with interruptible bidding that trades reliability for price.
Developer Tools - 04Modal
Serverless cloud running Python functions, GPU workloads, and scheduled jobs from a decorator, with no infra to manage.
Developer Tools - 05Baseten
Managed serving stack for deploying and autoscaling machine learning models on dedicated GPUs.
Developer Tools - 06CoreWeave
Kubernetes-native GPU cloud built for large training and inference fleets, with on-demand and reserved capacity.
Developer Tools
A curated selection in editorial order. Use the fit and evidence to judge it for your task. Something missing?
What to look for
- 01Serverless or reserved? Duty cycle decides this, and the cost difference is large in both directions.
- 02How long is a cold start? Loading a large model can dominate latency for infrequent requests.
- 03Is the memory enough for your model at your precision, with room for the batch?
Common questions
- Is it cheaper to run inference myself or use an API?
- Hosted APIs are cheaper until utilization is high and steady. Self-hosting pays off with sustained volume, strict data-residency requirements, or a fine-tuned model no API offers.
- What GPU memory do I need?
- Roughly two bytes per parameter at 16-bit, half that at 8-bit, a quarter at 4-bit, plus overhead for context and batch. A 7B model at 4-bit fits comfortably in 8GB.
More in Ship and operate
- Interface with your agentsCLI harnesses, IDEs, control planes, desktop apps, multiplexers, and terminals for steering coding agents.
- Deploy an applicationPublish previews and production builds without managing servers.
- Add a backendCombine databases, storage, APIs, and server-side functions.
- Add authenticationImplement accounts, sessions, identity providers, and authorization.
- Accept paymentsAdd subscriptions, checkout, billing, and payment infrastructure.
- Choose data infrastructureCompare databases, object storage, vector search, ORMs, and managed data services.
- Monitor product and usageCompare error monitoring, observability, product analytics, and web analytics.
- Automate deliveryBuild, test, preview, and release changes through CI/CD services.
- Add email and messagingSend transactional email, notifications, chat, and product messages.
- Protect agent credentialsCompare vaults, short-lived credentials, and egress proxies that keep secrets out of agent context.