// the find
cortexlabs/cortex
Production infrastructure for machine learning at scale
Cortex is Kubernetes infrastructure for deploying ML models to production on AWS — it wraps EKS with autoscaling, spot-instance handling, and three workload types (realtime, async, batch) so you get an endpoint instead of hand-building a serving stack. Aimed at teams already committed to AWS/EKS who want inference infra without assembling Istio, Prometheus, and autoscalers themselves.
Distinguishes autoscaling triggers per workload shape — in-flight request count for realtime, queue depth for async — instead of a generic HPA bolted onto everything. Spot instance support with automated on-demand fallback is a real operational headache it solves for you. Ships a CRD-based batch job controller (proper kubebuilder scaffolding, not just YAML templating) plus pre-built Grafana dashboards and CloudWatch log streaming out of the box.
The README states it's no longer maintained by its original authors and the last push was June 2024 — you'd be running EKS, Istio, and Prometheus glue code with no one patching CVEs in any of it. Hard AWS lock-in (IAM, VPC, EKS baked in throughout) with zero path to other clouds. The dependency stack is large (Istio proxy, cluster-autoscaler, custom device plugins, kube-rbac-proxy) so there's a lot that can break with nobody upstream to fix it, and the deployment story for canary/shadow rollouts is thin next to purpose-built serving tools like KServe or Seldon.