// the find
lllyasviel/Omost
Your image is almost there!
Omost fine-tunes small LLMs (Llama3-8B, Phi3-mini) to output Python code that builds a structured `Canvas` object describing image composition — global scene plus per-region bounding boxes, colors, and depth — which a bundled attention-manipulation renderer then turns into an actual diffusion-generated image. It's for people experimenting with LLM-driven layout control over image generation rather than plain text-to-image prompting.
The sub-prompt design is genuinely well thought out: every description is capped under 75 CLIP tokens and semantically self-contained, so a greedy bag-merge can pack them for encoding without ever truncating mid-concept. The spatial representation (9 locations x 9 offsets x 9 areas = 729 boxes) is a pragmatic, well-justified choice over raw pixel coordinates — the README explains they tried numeric coords and embeddings first and both trained poorly on Llama3/Phi3. The baseline renderer manipulates attention scores directly (per Dense Diffusion) rather than just masking Q/K/V, which is a real step up in regional fidelity without adding trained parameters or style drift.
No commits since July 2024 — over a year stale, and it's built on Llama3/Phi3-8B, which are aging base models with no sign of a refresh. It's a demo, not a library: two Gradio scripts and a `lib_omost` folder, no packaging, no tests, and the setup instructions (conda, specific torch/cu121 wheel, bitsandbytes) will break or need adjustment on newer CUDA/driver stacks. The `omost-dolphin-2.9-llama3-8b` variant is explicitly unfiltered/uncensored, which means anyone using it in a service has to bolt on their own safety layer — a real liability that isn't optional tooling, it's a deployment requirement the repo just punts on.