Medium potentialLow license · Apache-2.0Score 9/10

WASTE Cloud

Run trillion‑parameter LLMs on local hardware

803CBuild: 5-7 dayssqliteai/waste

A hosted inference service that lets you run the full Kimi K3 model on your own NVMe‑powered servers, with an OpenAI‑compatible API. It eliminates cloud inference costs and reduces latency for real‑time applications.

Tech stack

Docker + WASTE + FastAPI + Stripe + DigitalOcean Droplet with NVMe

BYOK model

Kimi K3 model weights (1.42 TB) and NVMe storage

Freemium model

Free tier: 100k tokens/month, 1 instance, 64 GB RAM, 1 TB NVMe. Pro tier: unlimited tokens, 2 instances, GPU optional, priority support.

Deployment complexity

Medium — The engine is dependency‑free and can be containerized, but requires large NVMe storage and 64 GB RAM, making hosting non‑trivial.

SEO keywords
on‑prem LLM inferencerun Kimi K3 locallytrillion parameter model inferenceself‑hosted LLM servicelow latency LLM inference
Why it qualifies
  • Large model support
  • No external dependencies
  • OpenAI API compatibility
  • High demand for on‑prem inference
Build this
ShareXLinkedInHN

Related analyses