finds.dev← search

// the find

Trinkle23897/learning-beyond-gradients

★ 630 · Python · updated May 2026

Heuristic Learning Blog Post

Public artifact repo for the blog post 'Learning Beyond Gradients'. Judging by what it ships, the post argues that code-level heuristics refined by search can compete with gradient-based RL on several tasks. It holds the policy scripts, raw trial logs, summaries and figures for Atari (Pong, Breakout, Montezuma, Atari57), MuJoCo (Ant, HalfCheetah) and VizDoom, plus the renderer that builds the bilingual article page. It's useful for anyone checking the article's numbers or rerunning the heuristic loop, but it is not a library.

- Raw per-trial records (`*_trials.jsonl`) sit next to the summary CSVs, so the article's numbers can be traced back to individual runs rather than a single headline score.

- The Montezuma folder keeps the search variants that were tried (archive search, state-graph search, ALE state search) with their own trial logs, which shows what didn't work as well as what did.

- `heuristic_learning/` is a separate package with tests for policies, search and the ledger, and it runs across several Gym-style environments, so the method can be exercised beyond the article's own tasks.

- Setup is left to the reader. The README says the article commands assume EnvPool 1.1.1 and the Atari and MuJoCo runtime are already installed, and its only install step covers the renderer.

- Most headline results come from hand-shaped per-game scripts with constants and macro sequences baked in (the Montezuma macros JSON is a clear example), so transfer to other tasks isn't shown here.

- The README doesn't say whether reported scores come from seeds unseen during search, and some scripts select the best trial, so the numbers may be optimistic next to a fresh evaluation.

- The root directory has unnamed `ig_*.png` files and many large MP4s committed directly, which makes the repo heavy to clone and hard to browse.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →