// the find
chime/terraform-aws-alternat
High availability implementation of AWS NAT instances.
alterNAT is a Terraform module that replaces AWS NAT Gateways with self-managed NAT instances, one Auto Scaling Group per availability zone, with a standby NAT Gateway in each zone for failover. It suits teams whose NAT Gateway data processing charges are large enough to justify operating EC2 instances, which the README puts at roughly 10TB a month. At low traffic the standby gateways still carry their hourly charge, and there is little data processing fee left to save.
Failover leans on managed parts. The standby NAT Gateway is AWS's responsibility, and the replace-route Lambda is the only custom control piece: it flips the route table and runs a connectivity check every minute from each private subnet. There is no second instance to keep healthy. Patching is replacement. Max instance lifetime is on by default at 14 days, and each boot pulls the latest vanilla Amazon Linux 2023 AMI, so there is no AMI pipeline to maintain. The failure modes are written down. The connection table loss, the restore flapping case, and the control-plane edge case each get a section, and the cost math shows the threshold where the module stops being worth it. Most modules of this kind skip that part.
Every instance replacement resets NAT state. The connection table lives on the instance, so each rotation breaks established flows through that zone, which at the 14-day default means once per zone per fortnight. Clients have to retry. The README's answer for long-lived transfers is to disable max lifetime, which gives up the patching story. The restore path can flap. If a security group rule is missing, curl from the instance over SSM still succeeds, the route moves back to the NAT instance, real traffic fails, and the Lambda moves it back to the NAT Gateway on the next run. The restore feature is off by default, but the README describes this loop without offering a guard against it. Failover depends on the EC2 API, which the Lambda reaches through an interface VPC endpoint. The README's edge case, where the control plane is down during an instance failure, still needs a human to change routes by hand. Operating it means more work than the pitch suggests. There is no published image, so you build and push the Lambda container yourself, and AWS caps internet-bound bandwidth at 5 Gbps for current-generation instances under 32 vCPUs, so instance sizing takes real measurement.