Scale to zero by design
Release accelerators when traffic disappears. A lightweight gateway keeps the endpoint available and signals KEDA when the next request arrives.
Hearth is a minimal control plane for bursty LLM workloads on private Kubernetes clusters—declarative model lifecycle, cold-start-aware traffic, and reusable runtime profiles without a fleet-scale platform.
$ kubectl get llmservice
demand preserved
backend scheduled
model ready · streaming
Why Hearth
Scaling a Deployment to zero is the easy part. Hearth concentrates on what happens when demand returns.
Release accelerators when traffic disappears. A lightweight gateway keeps the endpoint available and signals KEDA when the next request arrives.
Bound admission, preserve activation demand, emit streaming heartbeats, wait for model readiness, and drain in-flight requests safely.
Run existing vLLM images across NVIDIA and Ascend, use vendor device plugins, and opt into Volcano or observability without making them core dependencies.
Composable by boundary
Application owners declare serving intent. Cluster administrators publish reusable runtime profiles. Hearth translates both into the workloads and lifecycle resources your cluster already understands.
One cluster, two serving policies
Fleet routing, cache-aware scheduling, disaggregation, and continuously ready models.
A small declarative control plane for occasional models that should release their accelerators.
Try Hearth
Install the chart, select a hardware profile, and watch the backend wake from zero.