finds.dev← search

// the find

cortexlabs/cortex

★ 8,010 · Go · Apache-2.0 · updated Jun 2024

Production infrastructure for machine learning at scale

Cortex is Kubernetes infrastructure for deploying ML models to production on AWS — it wraps EKS with autoscaling, spot-instance handling, and three workload types (realtime, async, batch) so you get an endpoint instead of hand-building a serving stack. Aimed at teams already committed to AWS/EKS who want inference infra without assembling Istio, Prometheus, and autoscalers themselves.

Distinguishes autoscaling triggers per workload shape — in-flight request count for realtime, queue depth for async — instead of a generic HPA bolted onto everything. Spot instance support with automated on-demand fallback is a real operational headache it solves for you. Ships a CRD-based batch job controller (proper kubebuilder scaffolding, not just YAML templating) plus pre-built Grafana dashboards and CloudWatch log streaming out of the box.

The README states it's no longer maintained by its original authors and the last push was June 2024 — you'd be running EKS, Istio, and Prometheus glue code with no one patching CVEs in any of it. Hard AWS lock-in (IAM, VPC, EKS baked in throughout) with zero path to other clouds. The dependency stack is large (Istio proxy, cluster-autoscaler, custom device plugins, kube-rbac-proxy) so there's a lot that can break with nobody upstream to fix it, and the deployment story for canary/shadow rollouts is thin next to purpose-built serving tools like KServe or Seldon.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →