Managed LiteLLM VPS hosting
LiteLLM, live in minutes.
On a server that is yours alone.
One API gateway for 100+ LLM providers: unified OpenAI-format calls, spend tracking, rate limits, and failover, self-hosted.
- 500% SLA
- 30-day money-back
- Free Daily Snapshots
- Firewall & IDS
What it is
LiteLLM is the traffic controller for AI APIs: one self-hosted gateway that speaks OpenAI format and routes to 100+ providers, with per-key spend tracking, rate limits, retries, and failover. Point every app at one URL and manage the chaos in one place.
Minimum RAM
1 GB
Fits on SC-1GB; SC-2GB is the comfortable pick with room to stack.
What we install
Our playbook deploys the LiteLLM proxy with a PostgreSQL-backed key store, HTTPS issued and forced, and an admin key generated for you. Add provider keys, mint virtual keys per app or teammate, done.
Every plan includes
Instant hostname with TLS, database created and wired, daily snapshots, firewall and IDS. Add Fully Managed anytime and our engineers run the whole server for you.
Pick your server
Choose a plan for LiteLLM.
LiteLLM needs at least 1 GB of RAM. Every plan is your own private VPS, and the same server runs as many apps as fit.
SC-1GB
- 1 vCPU
- 1 GB RAM
- 25 GB NVMe
- 2.5 TB transfer
- Full Root Access
- Non-Oversold Cores
- 100% NVMe Storage
- Network Firewall/IDS
- Free Daily Snapshots
- DDoS-Protected Network
- Dual Backups (Optional)
- Fully Managed (Optional)
30-day money-back
SC-2GB
- 1 vCPU
- 2 GB RAM
- 50 GB NVMe
- 5 TB transfer
- Full Root Access
- Non-Oversold Cores
- 100% NVMe Storage
- Network Firewall/IDS
- Free Daily Snapshots
- DDoS-Protected Network
- Dual Backups (Optional)
- Fully Managed (Optional)
30-day money-back
SC-4GB
- 2 vCPU
- 4 GB RAM
- 100 GB NVMe
- 10 TB transfer
- Full Root Access
- Non-Oversold Cores
- 100% NVMe Storage
- Network Firewall/IDS
- Free Daily Snapshots
- DDoS-Protected Network
- Dual Backups (Optional)
- Fully Managed (Optional)
30-day money-back
SC-6GB
- 3 vCPU
- 6 GB RAM
- 150 GB NVMe
- 15 TB transfer
- Full Root Access
- Non-Oversold Cores
- 100% NVMe Storage
- Network Firewall/IDS
- Free Daily Snapshots
- DDoS-Protected Network
- Dual Backups (Optional)
- Fully Managed (Optional)
30-day money-back
LiteLLM installs itself on first boot with TLS, a hostname, and its database wired. Management and Dual Backups are optional addons in the cart.
Prefer to self-host? The exact LiteLLM image we deploy is public and free to pull, with a full production guide.
Why here
Why run LiteLLM on RemarkableCloud.
Your own private VPS
Not a container on a shared box. A real virtual server with dedicated resources and your own IP, and nobody else on it.
Stack apps, one price
Add as many apps as fit on your server from the client area, in one click each. No per-app charge, ever.
A 500% SLA that pays
1 hour down = 5 hours credited back, automatically, from the first minute. While others credit after thresholds, we credit from minute one.
Non-oversold hardware
Enterprise dedicated infrastructure, 100% NVMe, resources never sold twice. Your LiteLLM performs the same on the worst day as the best.
Humans since 2001
Real engineers answer, at any hour. Twenty-five years of running production servers, and no chatbot wall between you and them.
Free migration, done by us
Moving an existing LiteLLM site? We migrate it, verify it, and hand it back running, at no cost. Your old host stays live until you approve.
Built for
What people run on a LiteLLM VPS here.
One API key for every model
Your apps call one OpenAI-compatible endpoint; LiteLLM routes to any provider behind it.
Spend control across providers
Budgets, rate limits and usage logs per key, so AI costs stop being a mystery.
Provider failover
When one API degrades, routing rules fail over to another without touching app code.
Team key management
Hand each project a virtual key with its own limits instead of sharing raw provider keys.
The details
LiteLLM here, in plain words.
Why put a gateway in front of AI providers?
Because five apps with five hardcoded API keys is how spend surprises and outages happen. LiteLLM gives each app a virtual key with its own budget and rate limit, tracks spend per key, and fails over to a second provider when the first one has a bad day.
What does OpenAI-compatible actually mean?
Any tool that can talk to OpenAI's API can talk to your LiteLLM URL unchanged, while the gateway translates to Anthropic, Google, Mistral, Ollama, or a hundred others behind the scenes. Switching providers becomes a config line, not a code change.
How much server does LiteLLM need?
It is the lightest app in our AI shelf: 1 GB runs the proxy for personal and small-team traffic. It almost always shares a server with the apps that call it, which is exactly what the cart meter is for.
Does LiteLLM see my prompts?
Requests pass through your gateway on your server on the way to the provider you chose; nothing is logged beyond what you enable. That is the point of self-hosting the control plane.
Everything around LiteLLM, in one honest list.
Everything unmarked ships with every plan. The starred items are optional paid addons, priced plainly, added in the cart or anytime later.
* Optional paid addon: part of Fully Managed or Dual Backups. One price per plan, addons priced plainly, no surprise renewals.
Stack it
LiteLLM runs well with.
Same server, no extra bill. These are the companions our team installs next to LiteLLM most often.
LibreChat
+4 GBA private ChatGPT-style interface for every major model provider. One clean UI, your API keys, your conversation history on your server.
Flowise
+2 GBDrag-and-drop builder for LLM apps and agents. Prototype chatbots and pipelines visually, then serve them as APIs from your own server.
PostgreSQL
+1 GBThe engineers’ database. Installed, tuned for your RAM, and backed up twice offsite.
How it works
Seven systems working so you do not.
Every managed server runs the same stack underneath. Pick a system to see what it actually does.
Issues end before you hear about them.
Every minute, we check more than 200 parameters of your server and its services: web server, database, mail, disk, memory, and the network path to it. The moment anything drifts out of range, our support staff is alerted, investigates, and fixes it.
This is internal monitoring, deeper than any public uptime check: it watches the inside of your server, not just whether it answers. Most incidents are found and fixed by our staff before anyone outside ever notices them.
- 200+ parameters checked every minute
- Alerts reach engineers, not ticket queues
- Inside-the-server checks, not just ping
- Humans respond in minutes, any hour
What customers say
More reviews: Trustpilot“Excellent webhosting company. I really love all their products and services. Thank you so much.”
“The tech support is excellent, and they offer very competitive pricing. They also migrate my cPanel VPS to a DirectAdmin VPS without noticeable downtime.”
“I've been hosting several websites with them, and I can honestly say their support is outstanding. They are incredibly patient, knowledgeable, and genuinely committed to solving problems, even the most difficult and unusual ones.”
Frequently asked questions
Who should run LiteLLM?
Anyone with more than one AI-consuming app or more than one teammate holding raw provider keys. It is the natural companion to LibreChat, Flowise, and OpenClaw on the same server, all $0 here.
Does the gateway add latency?
Single-digit milliseconds on the same server, which disappears next to model inference time measured in seconds. Failover and retries typically save far more time than the hop costs.
What does LiteLLM hosting cost?
The app is $0; a 1 GB server starts at $4.17/mo on annual billing and runs it with room to spare, management and backups optional. Provider usage bills with each provider through your own keys.
Does it work with my existing OpenAI-client code?
Yes, that is the point: point the base URL at your LiteLLM and existing OpenAI SDK code runs unchanged while gaining routing, budgets and logs.
Can it route to local models too?
Yes, Ollama and other local endpoints register as providers alongside the commercial APIs, so the same key can serve private and hosted models by rule.
Is this LiteLLM on a VPS or on shared hosting?
A VPS. LiteLLM runs on your own private virtual server with dedicated, non-oversold resources and its own IP. You get real isolation and performance, with the setup, TLS and hardening handled for you.
Can I run other apps on the same server?
Yes. Add any other catalog app to this server from your client area, one click each, no per-app charge. How much fits is a question of RAM, and resizing is minutes.
Do you migrate my existing LiteLLM?
Yes, migration help is free. Open a ticket with access to the current setup and we move it, verify it, and hand it back running. Nothing switches until you approve.
Can I upgrade or downgrade the plan later?
Anytime. Plans resize from the client area in minutes, annual billing keeps its flat 17% discount at every size, and there is never a fee to change tiers.
Sizing an app stack? The VPS sizing calculator recommends a plan from your apps and traffic, with the math shown.
Shared CPU servers
The right home for LiteLLM and most stacks: 3.0+ GHz vCPU, from $4.17/mo.
See Shared CPU →Dedicated CPU servers
For CPU-hungry stacks and busy databases: cores that are physically yours.
See Dedicated CPU →All plans and pricing
Every plan, both families, one honest price with everything included.
See pricing →Your server runs. You sleep.
Fully managed hosting from people who have been doing this since 2001.