GPU cloud, minus the enterprise theater

GPU instances and model hosting for people who do not want to run their own metal.

Rent a GPU by the hour, or hand us a model and we serve it behind an OpenAI-compatible API. Instances from $0.19 an hour. No contracts, no minimum spend, billing per second.

2.1M GPU-hours served since 2023 · no credit card required to start

The GPUs4All console: a table of GPU instances with live status, cost, and GPU usage, plus spending cards.
console.gpus4all.dev / instances
Trusted by teams shipping models in production
codecentric Giskard Lakera deepset Konfuzio Weaviate Jina AI

Two products, one console

Rent a raw GPU, host a model, or both. Either way it is tracked in one console and billed per second.

GPU instances

$0.19 /hour

from, RTX 4090

  • Root access, persistent disk, public IP
  • Your image or one of ours
  • Up and online in a few minutes
  • Stop billing the moment you delete

Instance pricing →

Model hosting

$29 /month

plus per-token usage

  • OpenAI-compatible endpoint
  • You pick the GPU, we run it
  • Autoscaling with per-second compute billing
  • Most Llama, Qwen, Mistral checkpoints deploy fast

API overview →

How it works

Three steps, none of them involve talking to a salesperson.

1. Pick a GPU

4090 for a few bucks a day, H100 for real training runs. Create it, it is booted within minutes, you get an IP.

2. Point it at your work

SSH in and run your container, or upload a model and our hosting serves it for you. Either way, your code stays yours.

3. Pay per second

Billing stops the moment the instance is gone. A weekend batch job costs a few dollars, and you keep your storage.

Who uses us

Evaluation labs, product companies serving a model to their app, researchers who need a GPU for a weekend. A few bigger teams run batch jobs with us because per-second billing makes idle time nearly free.

"We moved a 7B to their hosted API and the bill is smaller than the Kubernetes cluster it replaced. The spending cap is what actually sold us."

Marco Villanueva, ML lead, OCTO Technology

"Rented an H100 on a Friday, fine-tuned over the weekend, paid about $150, deleted it. Nobody else makes that this easy."

Priya Raman, research engineer, Seldon

By the numbers

2.1M

GPU-hours served since 2023

~300

models hosted on our inference service

3

regions: Frankfurt, Ashburn, Singapore

4

people on the team, and someone is always on call

99.2%

uptime over the last 90 days

$0.19

cheapest hour of compute you can rent

Start free with $10 of compute

No credit card, no contract, no call. Create an account and the credit is there.

Create your account