finds.dev

public digest · 5 picks

This week: nginx internals, JSON-shaped APIs, and a reality check on 'AI-native'

Five repos this week span from foundational infrastructure to teaching demos to things trying to bolt AI onto old problems. Some of it is code you'll run in production, some of it is code worth reading to see how a pattern should be done.

As always, being on this list means we think it's worth your time — the caveats below are things to know going in, not reasons to skip.

// pick 1 of 5

loiane/crud-angular-spring

A two-repo teaching project pairing a Spring Boot 4 / Java 25 REST API with an Angular v22 front end, built around a single Course-has-many-Lessons relationship. It's aimed at developers who already know the basics of both stacks and want to see a fuller-than-usual example — proper DTOs, validation, error handling, and tests on both sides — rather than a bare hello-world CRUD.

This is what a CRUD tutorial looks like when someone actually cares about the boring parts. Testcontainers against real MySQL, Pact contracts between the Angular and Spring sides, ArchUnit enforcing layering, RFC 7807 error responses — most demo repos stop at 'it compiles,' this one has a real test pyramid.

The Angular side is current too: standalone components, signals, zoneless change detection, signal forms. If you've been meaning to see what modern Angular looks like beyond blog posts, this is a working reference.

What to know going in: the README's own 'not implemented' list includes auth, caching, and rate limiting, so 'close to production-ready' is generous marketing on an otherwise honest project. It's also MySQL-specific in practice despite claims otherwise, and the domain model is deliberately thin — one has-many relationship — so you're learning the pattern, not stress-testing it.

View on GitHub → Our full take →

// pick 2 of 5

datopian/portaljs

PortalJS is a Next.js framework from Datopian (the CKAN company) for building data-catalog sites — a home page, a searchable catalog, and per-dataset showcase pages, all reading through a swappable DataProvider so the backend can be flat files, CKAN, GitHub, or a Parquet+DuckDB lakehouse. Its hook is a set of Claude Code slash commands (/portaljs-new-portal, /portaljs-add-dataset, etc.) that scaffold and wire up the portal for you. Best for teams standing up a public data catalog who don't want to build search/showcase/metadata pages from scratch.

A Next.js framework for data catalogs with a genuinely swappable backend — the DataProvider abstraction is exercised against CKAN, GitHub, and Frictionless in the examples folder, not just described in a diagram. There's also a real escape hatch: a tiged command gets you the plain template with zero AI dependency, so the 'AI-native' framing doesn't lock you out if you don't want it.

The storage defaults (git + R2 + Parquet + DuckDB) are a sane middle ground for teams who don't want to either wrangle loose CSVs or stand up a warehouse they don't need.

Heads-up: the AI hook is Claude Code slash commands specifically — if you're on anything else, you're just using the plain template, and most of the pitch doesn't apply to you. Scope has also crept past 'frontend framework' — there's a whole Cloudflare backend and a hosted deploy target baked in — so 'no lock-in' is truer of the raw frontend code than of the full workflow.

View on GitHub → Our full take →

// pick 3 of 5

APIJSON/APIJSON

APIJSON is a Java library where the client sends a JSON document describing the shape of data it wants — including joins across tables — and the server parses that into SQL and returns matching JSON, instead of you writing a REST endpoint per use case. It's aimed at backend teams who are tired of hand-rolling CRUD endpoints and want the client to shape its own response.

The core idea here is worth stealing even if you never use the library: request and response share the same JSON shape, so a client can ask for a user plus their last 5 orders in one call without a bespoke endpoint or standing up GraphQL. Database support is unusually wide (MySQL, Postgres, ClickHouse, Elasticsearch, Milvus, and more), and this isn't a toy — 18k stars, active commits, named production users, and public security audits from Tencent and Ant Group.

What to know going in: letting the client construct arbitrary joins is a large attack surface by design, and the whole security model rests on getting the Access/Verifier config exactly right — get it wrong once and you've built an IDOR generator. The README is mostly social proof; the real documentation lives in a separate Document.md. And don't point production traffic at their public demo server, tempting as the 'fastest path' framing makes it sound.

View on GitHub → Our full take →

// pick 4 of 5

nginx/nginx

nginx is the C-based web server, reverse proxy, load balancer, and mail proxy that quietly runs a huge chunk of the internet's HTTP traffic, and this repo is now growing HTTP/3/QUIC support directly in core. It's for anyone standing up production HTTP infrastructure, from a single box terminating TLS in front of an app server to a CDN edge node.

The web server running a huge share of the internet's HTTP traffic, still under active development, now growing native HTTP/3/QUIC support in core rather than as a bolt-on module. The event-driven worker model is the reason it displaced thread-per-connection servers for high-concurrency workloads, and the config DSL genuinely handles load balancing, rate limiting, and caching without reaching for extra software.

Worth reading the source even if you'll never contribute — the module system and memory pool design are a masterclass in doing a lot with a small, disciplined codebase.

Heads-up: contributing means learning nginx's own conventions (ngx_pool) on top of C itself, and most commits still come from F5 employees. The dynamic module ABI isn't stable across versions, so third-party modules often need rebuilding after a minor bump. And some things tutorials assume nginx does — active health checks, session persistence — actually live behind the closed-source NGINX Plus.

View on GitHub → Our full take →

// pick 5 of 5

m3dev/gokart

gokart is m3's opinionated wrapper around Luigi that adds reproducibility guarantees to ML pipelines — every task's output, parameters, module versions, and even random seed get hashed and cached to a pkl file, so reruns are automatic when inputs change and skipped when they don't. It's aimed at teams running batch ML pipelines in Python who want Makefile-style caching without hand-rolling it, and who are fine writing config as Python classes rather than YAML.

A Luigi wrapper that adds real reproducibility to ML pipelines: every task's output, parameters, and even random seed get hashed and cached, so changing one hyperparameter only reruns what's actually downstream of it. The mypy plugin with generic task types is the standout feature — it catches wiring mistakes between tasks at typecheck time instead of three hours into a batch job, which is rare in this space.

Three years of production use at m3 plus a documented competition win is more real validation than most pipeline frameworks this size can show.

What to know going in: it inherits Luigi's single-process central scheduler, so there's no distributed execution story here. Experiment tracking and DAG visualization are both punted to a separate repo or nonexistent, and every intermediate gets serialized to disk/S3 as pkl by design — fine for typical pipelines, a real bottleneck if your intermediates are large tensors or dataframes.

View on GitHub → Our full take →

If you want this in your inbox instead of hunting for it on the site, the email signup is at the bottom of the page.

Get this in your inbox →