Managed Ollama VPS hosting
Ollama, live in minutes.
On a server that is yours alone.
Run open LLMs on your own server. Honest requirement: real models want real RAM, and the meter will route you to the right plan.
- 500% SLA
- 30-day money-back
- Free Daily Snapshots
- Firewall & IDS
What it is
Ollama runs large language models on your own server: Llama, Mistral, Qwen, and the open-model world behind one clean local API, with your prompts never leaving the machine.
Minimum RAM
16 GB
Fits on SC-16GB; SC-48GB is the comfortable pick with room to stack.
What we install
The playbook deploys the current Ollama with the API bound privately by default, your chosen starter models pulled, memory sized to the plan, and the service supervised. Open WebUI beside it is one more checkbox.
Every plan includes
Instant hostname with TLS, database created and wired, daily snapshots, firewall and IDS. Add Fully Managed anytime and our engineers run the whole server for you.
Pick your server
Choose a plan for Ollama.
Ollama needs at least 16 GB of RAM. Every plan is your own private VPS, and the same server runs as many apps as fit.
SC-16GB
- 6 vCPU
- 16 GB RAM
- 400 GB NVMe
- 40 TB transfer
- Full Root Access
- Non-Oversold Cores
- 100% NVMe Storage
- Network Firewall/IDS
- Free Daily Snapshots
- DDoS-Protected Network
- Dual Backups (Optional)
- Fully Managed (Optional)
30-day money-back
SC-24GB
- 8 vCPU
- 24 GB RAM
- 600 GB NVMe
- 60 TB transfer
- Full Root Access
- Non-Oversold Cores
- 100% NVMe Storage
- Network Firewall/IDS
- Free Daily Snapshots
- DDoS-Protected Network
- Dual Backups (Optional)
- Fully Managed (Optional)
30-day money-back
SC-48GB
- 12 vCPU
- 48 GB RAM
- 1200 GB NVMe
- 120 TB transfer
- Full Root Access
- Non-Oversold Cores
- 100% NVMe Storage
- Network Firewall/IDS
- Free Daily Snapshots
- DDoS-Protected Network
- Dual Backups (Optional)
- Fully Managed (Optional)
30-day money-back
SC-64GB
- 16 vCPU
- 64 GB RAM
- 1600 GB NVMe
- 160 TB transfer
- Full Root Access
- Non-Oversold Cores
- 100% NVMe Storage
- Network Firewall/IDS
- Free Daily Snapshots
- DDoS-Protected Network
- Dual Backups (Optional)
- Fully Managed (Optional)
30-day money-back
Ollama installs itself on first boot with TLS, a hostname, and its database wired. Management and Dual Backups are optional addons in the cart.
Why here
Why run Ollama on RemarkableCloud.
Your own private VPS
Not a container on a shared box. A real virtual server with dedicated resources and your own IP, and nobody else on it.
Stack apps, one price
Add as many apps as fit on your server from the client area, in one click each. No per-app charge, ever.
A 500% SLA that pays
1 hour down = 5 hours credited back, automatically, from the first minute. While others credit after thresholds, we credit from minute one.
Non-oversold hardware
Enterprise dedicated infrastructure, 100% NVMe, resources never sold twice. Your Ollama performs the same on the worst day as the best.
Humans since 2001
Real engineers answer, at any hour. Twenty-five years of running production servers, and no chatbot wall between you and them.
Free migration, done by us
Moving an existing Ollama site? We migrate it, verify it, and hand it back running, at no cost. Your old host stays live until you approve.
Built for
What people run on a Ollama VPS here.
Private LLMs for real work
Llama, Mistral, Qwen and friends answering on your server. Prompts and outputs never leave it.
An API for your applications
OpenAI-compatible endpoints mean your app code barely changes; only the bill and the privacy do.
Confidential-document workflows
Summarize contracts and internal docs with a model that cannot leak them to a provider.
Model experimentation
Pull, compare and swap models freely. The meter is your server, which does not meter.
The details
Ollama here, in plain words.
What does managed Ollama hosting include?
A working local-model server sized honestly, kept updated, and secured by default, at flat pricing with no per-token anything. The API is yours; the meter does not exist.
What can a CPU server realistically run?
The honest answer: small and mid-size models. 7B-class models run acceptably for chat and drafting on 16 GB; quantized 13B fits with patience; large models want GPUs we do not sell, and we will say so rather than sell you disappointment. For summaries, extraction, and private chat, CPU-served 7B is genuinely useful.
Why run models locally at all?
Privacy and cost shape. Client documents, contracts, and internal data can be processed without leaving your server, which turns entire compliance conversations into a shrug. And a flat server price beats per-token billing the moment usage becomes routine.
What pairs with Ollama?
Open WebUI for a ChatGPT-style interface on top, and n8n for automation that calls the models: summarize inbound email, draft replies, classify tickets, all over localhost, all private. The Private AI kit bundles exactly this.
Everything around Ollama, in one honest list.
Everything unmarked ships with every plan. The starred items are optional paid addons, priced plainly, added in the cart or anytime later.
* Optional paid addon: part of Fully Managed or Dual Backups. One price per plan, addons priced plainly, no surprise renewals.
Stack it
Ollama runs well with.
Same server, no extra bill. These are the companions our team installs next to Ollama most often.
Open WebUI
+2 GBThe chat interface for Ollama. Together they are a private ChatGPT on hardware nobody shares.
n8n
+2 GBSelf-hosted workflow automation: connect your apps and let the flows run 24/7 on a server that is monitored for you.
How it works
Seven systems working so you do not.
Every managed server runs the same stack underneath. Pick a system to see what it actually does.
Issues end before you hear about them.
Every minute, we check more than 200 parameters of your server and its services: web server, database, mail, disk, memory, and the network path to it. The moment anything drifts out of range, our support staff is alerted, investigates, and fixes it.
This is internal monitoring, deeper than any public uptime check: it watches the inside of your server, not just whether it answers. Most incidents are found and fixed by our staff before anyone outside ever notices them.
- 200+ parameters checked every minute
- Alerts reach engineers, not ticket queues
- Inside-the-server checks, not just ping
- Humans respond in minutes, any hour
What customers say
More reviews: Trustpilot“Excellent webhosting company. I really love all their products and services. Thank you so much.”
“The tech support is excellent, and they offer very competitive pricing. They also migrate my cPanel VPS to a DirectAdmin VPS without noticeable downtime.”
“I've been hosting several websites with them, and I can honestly say their support is outstanding. They are incredibly patient, knowledgeable, and genuinely committed to solving problems, even the most difficult and unusual ones.”
Frequently asked questions
How much RAM does Ollama need?
16 GB is the honest minimum for comfortable 7B-class use, and the meter enforces it. More RAM means bigger or less-quantized models.
Which models can I pull?
Anything in the Ollama library or importable as GGUF: Llama, Mistral, Qwen, Gemma, Phi, and the rest of the open world.
Is the API exposed?
Not by default: it binds privately, apps on the same server reach it over localhost, and remote access happens through WireGuard or a deliberate firewall grant.
Which models fit in my RAM?
Rule of thumb: quantized 7 to 8B models are comfortable at 16 GB, 13B-class wants more headroom. Larger plans open larger models; resizing is minutes when curiosity grows.
How fast is CPU inference, honestly?
Usable for chat and background jobs with small models: several tokens per second on dedicated cores, faster with more of them. It is not GPU speed, and for many private workloads it does not need to be.
Is this Ollama on a VPS or on shared hosting?
A VPS. Ollama runs on your own private virtual server with dedicated, non-oversold resources and its own IP. You get real isolation and performance, with the setup, TLS and hardening handled for you.
Can I run other apps on the same server?
Yes. Add any other catalog app to this server from your client area, one click each, no per-app charge. How much fits is a question of RAM, and resizing is minutes.
Do you migrate my existing Ollama?
Yes, migration help is free. Open a ticket with access to the current setup and we move it, verify it, and hand it back running. Nothing switches until you approve.
Can I upgrade or downgrade the plan later?
Anytime. Plans resize from the client area in minutes, annual billing keeps its flat 17% discount at every size, and there is never a fee to change tiers.
Sizing an app stack? The VPS sizing calculator recommends a plan from your apps and traffic, with the math shown.
Shared CPU servers
The right home for Ollama and most stacks: 3.0+ GHz vCPU, from $4.17/mo.
See Shared CPU →Dedicated CPU servers
For CPU-hungry stacks and busy databases: cores that are physically yours.
See Dedicated CPU →All plans and pricing
Every plan, both families, one honest price with everything included.
See pricing →Your server runs. You sleep.
Fully managed hosting from people who have been doing this since 2001.