finds.dev← search

// the find

davidzajac1/zillacode

★ 353 · Python · Apache-2.0 · updated Jun 2025

Open Source LeetCode for PySpark, Spark, Pandas and DBT/Snowflake

Zillacode is an open-source clone of LeetCode aimed at data engineers, letting you practice 50+ problems in PySpark, Spark, Pandas, and DBT/Snowflake instead of the usual algorithms-in-Python format. It's for people prepping for data engineering interviews who want to drill query and transformation logic rather than pointer arithmetic.

The problem domain is genuinely underserved — there's no other free platform drilling PySpark/DBT/Snowflake interview problems specifically. One-command Docker Compose startup for local use is a real convenience given the project spans 5 microservices (Flask, two Spark Lambda variants, a sandboxed Pandas exec service, React frontend). Leaving the original Terraform/Zappa/GitHub Actions IaC in the repo is a nice bonus for anyone who wants to study how a real serverless SaaS was deployed on AWS Lambda, even though the app no longer runs that way.

The Snowflake/DBT problems require you to bring your own Snowflake account and credentials, so a chunk of the advertised content doesn't work out of the box. The Scala Spark Lambda build via sbt takes ~15 minutes, which is a rough dev loop if you're touching that service. The DB Lambda's use of Python's exec() to run submitted Pandas code, while isolated into its own permission-less service, is the kind of design that needs constant vigilance — worth checking how submissions are actually sandboxed before trusting it with anything beyond a local toy deployment. Last push was June 2025, so it's not clear how actively new problems or fixes are still landing.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →