FastAPI VPS Requirements: Uvicorn Workers, RAM, HTTPS & Production Sizing
FastAPI VPS sizing guide covering Uvicorn workers, CPU/RAM, reverse proxy/HTTPS, database/cache memory, background tasks, logs, backups and scaling.
FastAPI VPS requirements depend on process count, request concurrency, application CPU time, database/cache memory and background work. A lightweight async API can serve significant traffic on a modest VPS, while image processing, ML inference, large JSON transforms or blocking third-party calls can require much more capacity.
FastAPI’s current deployment guidance covers Uvicorn worker processes and notes that multiple worker processes can take advantage of multiple CPU cores. More workers also consume more memory, so worker count should be tested rather than copied blindly.
A production FastAPI stack
- FastAPI application;
- Uvicorn workers (or a process manager/container strategy);
- reverse proxy/load balancer where appropriate;
- HTTPS/TLS termination;
- PostgreSQL/MySQL/MongoDB if the app uses a database;
- Redis/cache/queue if required;
- background workers;
- monitoring/logging.
How many workers?
Use enough workers to utilise available CPU and handle concurrency, but not so many that the VPS runs out of RAM or overloads the database. An I/O-bound API can handle many concurrent requests per worker, while a CPU-heavy endpoint can saturate a process quickly.
Measure memory per worker and p95/p99 latency while increasing concurrency. Stop adding workers when throughput stops improving or database/CPU saturation increases.
CPU requirements
Async does not make CPU work free. JSON serialization, cryptography, compression, image processing and ML inference can still consume CPU. If one worker spends most of its time doing CPU work, a high-frequency VPS may help. If many independent requests run in parallel, additional vCPU can help.
RAM requirements
Budget memory for every worker plus the OS and local services. If the database and Redis share the same VPS, reserve their memory explicitly. Containerised deployments also need memory limits that match the host capacity.
| Stack | Practical starting direction |
|---|---|
| Small API, external DB | 2–4 GB can be sufficient after testing |
| Production API + local DB/cache | 4–8 GB often provides safer headroom |
| Many workers / queues | 8–16 GB+ depending on process memory |
| CPU-heavy endpoints | Profile CPU; consider high-frequency compute |
Reverse proxy and HTTPS
Production deployments should use a proper HTTPS path, secure headers and a process/restart strategy. Whether TLS terminates at Nginx, a load balancer or another edge depends on the architecture. Restrict the application service to the intended network path rather than exposing every internal port publicly.
Database connection pooling
Worker count affects database connections. Ten workers each opening many connections can overwhelm a small database even when the API CPU is idle. Use appropriate pooling and observe active connections, query latency and timeouts.
Background work
Long-running or CPU-heavy jobs should not block normal API responses. Use background workers/queues where appropriate and size them independently. Export jobs, PDFs, media processing and ML work can be isolated on separate VPS nodes.
Logs and disk
Application/access logs can fill a VPS surprisingly quickly. Configure rotation and retention. Keep persistent user uploads and database backups outside ephemeral container layers.
Monitoring
- p50/p95/p99 request latency;
- requests per second;
- worker CPU and RSS;
- event-loop/request queue symptoms;
- database latency/connections;
- 5xx/timeouts;
- disk and network usage.
When to scale vertically vs horizontally
Increase VPS size when a single node is resource-constrained and the application is easy to scale vertically. Add more API nodes behind a load balancer when one machine becomes a failure or capacity bottleneck and the application is stateless enough to scale horizontally.
VPSWala fit
VPSWala’s KVM plans scale from small nodes through 128 GB RAM. For CPU-sensitive APIs, the 9950X range provides a high-frequency option. Choose from measurements rather than assuming every FastAPI service needs the same plan.
Bottom line
FastAPI itself is efficient, but production capacity depends on what each request does. Size Uvicorn workers, CPU and RAM together, protect the HTTPS path, control database connections, isolate heavy jobs and scale from observed latency and resource use.
Next step: compare VPSWala KVM VPS and 9950X VPS using your worker count, database choice and peak concurrency.
Frequently asked questions about FastAPI VPS hosting
Can FastAPI run on 1 GB RAM?
A very small API can, especially with an external database and one process, but 1 GB leaves little production headroom for the OS, logging and traffic spikes. Measure the real application because imported libraries and model/data objects can change memory use dramatically.
Should Gunicorn still be used with FastAPI?
FastAPI deployment patterns evolve with Uvicorn and container tooling. Follow the current FastAPI/Uvicorn guidance for your chosen deployment model rather than copying an old command from a tutorial. The core requirement is supervised production processes with restart/startup handling, HTTPS and measured worker count.
How do I know when to add another API node?
If vertical scaling no longer gives enough capacity, one VPS becomes a maintenance/failure bottleneck, or traffic must survive node loss, deploy multiple stateless API instances behind a load balancer. Externalise session/state so any node can serve the next request.
Sources
Not sure which size?
Send the stack, get a size.
Tell us the operating system, application stack, current traffic, database size and where it hurts today. You get a sizing recommendation, the matching plan and a price.
Related
Deploying Next.js SSR on a Linux VPS: Standalone Build, PM2 & Nginx Proxy
Learn how to deploy a server-rendered Next.js application on a Linux KVM VPS using standalone build artifacts, PM2 cluster management, and Nginx.
VPS Firewall Setup with UFW and iptables: Port Hardening & Lockout Prevention
A hands-on sysadmin guide to securing a Linux cloud VPS with UFW and iptables, implementing strict ingress filtering while preventing accidental connection loss.
VPS Hosting for Bhopal: Which VPSWala Node to Pick and How to Test It
A practical routing and workload sizing guide for developers and businesses in Bhopal to evaluate VPSWala cloud nodes and verify network path stability.