finds.dev← search

// the find

elastic/eland

★ 693 · Python · Apache-2.0 · updated Jun 2026

Python Client and Toolkit for DataFrames, Big Data, Machine Learning and ETL in Elasticsearch

Eland is Elastic's own Python client that exposes Elasticsearch-indexed data through a pandas-like DataFrame API, pushing filtering and aggregation down to Elasticsearch instead of pulling data into local memory. It also ships tooling to export trained scikit-learn/XGBoost/LightGBM models and Hugging Face transformer models into Elasticsearch's inference layer. It's for data scientists and ML engineers already on the Elastic stack who want to explore ES-backed datasets in a notebook or serve models through ES itself, not a general-purpose pandas replacement.

Real lazy evaluation — DataFrame operations translate to ES queries and aggregations, so you can explore an index far larger than your machine's RAM without hitting OOM. The model-import story is unusually broad for a single package: classic ML (sklearn, XGBoost, LightGBM) and NLP transformers via TorchScript export, both covered by one tool. It's maintained directly by Elastic with CI on Buildkite and an explicit, enforced compatibility matrix rather than a loose 'should work' claim. The `eland_import_hub_model` CLI gives a one-command path from a Hugging Face model ID to a deployed ES inference endpoint.

For the PyTorch/NLP path your eland minor version has to match your Elasticsearch cluster's minor version — a real constraint if you're not on the latest ES or run mixed-version clusters, and it's easy to hit without warning. The README flat-out states PyTorch model import can execute arbitrary code on your ES server, which is a security-relevant design wart, not just a footnote. The pandas-API coverage is partial by nature of being backed by ES aggregations, and the README doesn't document which operations silently fall back to materializing full result sets locally — that's the kind of thing you discover in production. The quick-start `--start` deployment flag is explicitly called out as giving bad throughput (one allocation, one thread), so the easy path for a first deploy is also the wrong one for anything beyond a toy test.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →