Jina embedding and reranker models with frontier-grade accuracy are now available with zero external calls
Elastic (NYSE: ESTC), the Search AI Company, today announced that Jina AI models are available for on-premises and air-gapped environments through Jina On-Prem. Designed for regulated industries, air-gapped environments, or organizations that want full control over data, cost, and performance, Jina On-Prem delivers enterprise-grade data extraction and semantic search with no internet connection or third-party AI services required.
Many organizations need AI systems that don't depend on a live connection to a third-party service. While self-hosted alternatives exist, they come with tradeoffs. Supporting multiple media types and languages typically requires assembling separate models, and licensed platforms cleared for air-gapped deployments are largely constrained to text and images. The most accurate open-source models also carry substantial compute requirements, meaning comprehensive coverage often demands running several large models simultaneously.
Jina On-Prem packages Jina AI's family of models that cover text, images, audio, and video in a single embedding space so they can now run entirely within a customer's own environment, making no calls to the outside once deployed. Data stays on-premises, and teams retain full control over cost, model access, and performance. Small Jina models run on a single 8GB GPU, at a fraction of the cost, while matching the accuracy of far larger models that need many times more GPU memory.
"Historically, teams running search and retrieval in regulated or disconnected environments have had to choose between capability and control," said Ajay Nair, general manager, Elasticsearch and Platform, Elastic. "The ability to run Jina models fully on-premises removes that compromise by giving them Jina AI’s high-performance reader, embedding and reranking models directly in their own environments, allowing them to build AI applications without depending on external AI services.”
Jina On-Prem installs with a single command and makes no outbound network calls once deployed: no license server, no telemetry or logging endpoint, and no connection to a model registry. The suite includes all 28 Jina AI models, including the jina-embeddings-v5-omni multimodal embedding model and jina-reranker-v3, and supports both CPU and GPU hardware with automatic GPU detection. Applications can reach it through standard API schemas so existing integrations work without rewriting code.
Jina On-Prem also serves as a drop-in replacement for models served through Elastic Inference Service (EIS), so air-gapped Elastic deployments can integrate it directly without changing how applications call their embedding and reranking models.
Availability
Jina On-Prem is available now for download via GitHub through an access token. Installation instructions are available on the Jina On-Prem Quick Start page, with a separate bundling guide for teams composing their own Docker container. To license Jina On-Prem, please contact Elastic Sales.
Additional Materials
About Elastic
Elastic (NYSE: ESTC), the Search AI Company, integrates its deep expertise in search technology with artificial intelligence to help everyone transform all of their data into answers, actions, and outcomes. Elastic's Search AI Platform — the foundation for its search, observability, and security solutions — is used by thousands of companies, including more than 50% of the Fortune 500. Learn more at elastic.co.
Elastic and associated marks are trademarks or registered trademarks of elasticsearch B.V. and its subsidiaries. All other company and product names may be trademarks of their respective owners.
View source version on businesswire.com: https://www.businesswire.com/news/home/20260727539736/en/
Contacts
Media Contact
Elastic PR
PR-team@elastic.co