High potentialLow license · Apache-2.0Score 10/10
KimiServe
Run 2.78T LLM on a single CPU
KimiServe is a hosted inference‑as‑a‑service platform that runs the Kimi K3 model on commodity hardware, delivering low‑latency text generation without GPUs. It offers a simple REST API, auto‑scaling, and cost‑effective pricing for small teams.
Tech stack
FastAPI + Uvicorn + Docker + PostgreSQL + Stripe + Auth0
BYOK model
Not required
Freemium model
Free tier: 1,000 tokens/month, 1 CPU instance, 8 GB RAM, 1.56 TB checkpoint, 1 request/second limit. Pro tier: unlimited tokens, 2 CPU instances, 16 GB RAM, 2× throughput, priority support, 5 requests/second.
Deployment complexity
Medium — Requires a 1.56 TB checkpoint and fast SSD, but can be packaged in a Docker image and deployed on a single server.
SEO keywords
CPU LLM inferenceKimi K3 hosted2.78T LLM on CPUlow-cost AI chatbotCPU-only large language model
Why it qualifies
- ✓High demand for low‑cost large‑model inference
- ✓Unique capability of running 2.78T model on CPU
- ✓Open-source permissive license allows commercial use
Build this