finds.dev← search

// the find

aryn-ai/sycamore

★ 607 · Python · Apache-2.0 · updated Aug 2026

🍁 Sycamore is an LLM-powered search and analytics platform for unstructured data.

Sycamore is a Python framework for building ETL pipelines over unstructured documents (PDFs, decks, transcripts) that chunks, enriches with LLM calls, and loads into vector stores or search engines. It's aimed at teams building RAG or search over messy document collections who want more structure than 'just chunk and embed' but don't want to hand-roll a Ray pipeline.

The DocSet abstraction (map/flatmap/filter/llm_map) gives a composable, lazy pipeline API in the Spark/Ray-Dataset style, with LLM-backed transforms for entity extraction, schema extraction, and table/summary handling baked in rather than left as an exercise. Connector coverage is genuinely broad - OpenSearch, Elasticsearch, Pinecone, DuckDB, Qdrant, Weaviate, Neo4j - installed as pip extras, so swapping a vector store is a config change, not a rewrite. Built on Ray, so it actually scales horizontally for large document batches instead of being a single-process script pretending to be a pipeline.

The headline numbers (6x chunking accuracy, 2x recall improvement) are the vendor's own benchmarks for the hosted Aryn DocParse API, stated with no linked methodology - take them as marketing, not measurement. The best partitioning quality is tied to signing up for that hosted service; running the partitioner locally is mentioned in passing with no indication of whether quality or speed hold up without it. Ray as a hard dependency for the core engine is a heavy price for teams that just want to batch-process a few thousand files locally - you inherit Ray's packaging and operational quirks even on a single machine. The repo is also a sprawling monorepo (separate apps for Jupyter, OpenSearch, a remote-processor-service, even a C++ timetrace tool) for a company's commercial product wrapped in an OSS core, which is worth knowing going in re: how much of the roadmap serves Aryn's hosted offering versus the standalone library.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →