finds.dev← search

// the find

lvy010/AI-wiki

★ 274 · Python · updated Oct 2026

AI Full Stack: Data, Algorithms, Models, Hardware, Architecture

A personal, mostly-Chinese knowledge wiki of papers, courses, and blog posts spanning LLM systems topics — attention kernels, speculative decoding, KV-cache/PD disaggregation, MoE, quantization, CUDA. It's aimed at people trying to build full-stack literacy in LLM infra, from algorithms down to hardware, not at people looking for a library to install.

The paper selection for inference systems is genuinely good and current — PagedAttention, Mooncake/DistServe/MemServe for prefill-decode disaggregation, Sarathi-Serve for chunked prefill batching — the kind of reading list you'd get from someone actually working on serving infra, not a generic 'awesome-llm' scrape. It's not pure links either: there's a working cs336 lab, a Rust tokenizer example with its own Cargo.toml, and a GRPO.py implementation under work/llm, so some of this has been exercised by hand.

A large fraction of the blog and course links are literally '(#)' placeholders with no URL — 'Harness Engineering', 'Learn Claude Code', several FlashAttention writeups all point nowhere. The repo is almost entirely Chinese prose with no English README or translation, which cuts off most of the audience this newsletter reaches. It's a link list with section headers, not annotated — no indication of why one paper is included over another, how they relate, or what order to read them in, so the 'full-stack architecture thinking' pitch in the preface isn't backed by structure in the content. The roadmap was targeted for September and the repo is already past that with no roadmap.md changes reflected in what's visible, so treat the organization as still in flux.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →