Chroma

Chroma

Chroma is an open-source data infrastructure for AI that provides fast, serverless, and scalable search capabilities. It supports vector search (dense and sparse), full-text search, regex search, and metadata filtering, all built on object storage for cost efficiency. Chroma allows storing embeddings with metadata and querying across text, images, and other modalities. Designed for developers and AI engineers, Chroma powers retrieval-augmented generation (RAG), code search, agentic search, and more. Features include automatic data tiering, zero-ops infrastructure, SOC 2 compliance, and enterprise options like BYOC in your VPC. With over 15 million monthly downloads and 27K GitHub stars, it is trusted by companies like Capital One and UnitedHealthcare. Licensed under Apache 2.0, Chroma can be self-hosted or used via Chroma Cloud for a managed experience.

Search
Visit website
Added on
Jul 13, 2026
Open-sourceApache 2.0Vector searchFull-text searchSparse searchMetadata filterObject-storageMulti-modal
Product information

Everything worth knowing about Chroma

The complete picture — from everyday features to the technical detail builders and IT folks dig for.

Core features

What it actually does

Vector Search
Semantic similarity search using dense embeddings.
Sparse Vector Search
Lexical search using BM25 and SPLADE for keyword-based retrieval.
Full-Text & Regex Search
Keyword and regular expression search over documents without requiring embeddings.
Metadata Filtering
Filter search results at query time by metadata conditions.
Collection Forking
Dataset versioning, A/B testing, and roll-outs with copy-on-write.
Technical capabilities
Object Storage Architecture
Built on S3/GCS with automatic query-aware data tiering for cost efficiency.
Zero-Ops Infrastructure
Auto-scales with usage, no manual tuning, serverless pricing model.
Must watch videos
Chroma Cloud - Now available via Stripe Projects - Allow agents to provision infrastructure
2:27
Tutorial2:27
Chroma Cloud - Now available via Stripe Projects - Allow agents to provision infrastructure
Chroma Cloud is now available on Stripe Projects. Stripe Projects allows agents to create accounts, connect billing and provision infrastructure on your behalf. https://projects.dev
Deep dive: Using Reranking to improve search experiences with Chroma Cloud
15:23
Review15:23
Deep dive: Using Reranking to improve search experiences with Chroma Cloud
Reranking is a critical step in making search more useful to users. There’s tons of user and application specific context that is useful to users that is lost if you simply return the results from search alone. Reranking considers signals and data from your application to place results that are more likely to satisfy the user’s search needs. Links - Learning to rank: https://en.wikipedia.org/wiki/Learning_to_rank - XGBoost: https://xgboost.readthedocs.io/en/release_3.2.0/ - LambdaMark: https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/MSR-TR-2010-82.pdf - LightGBM: https://github.com/lightgbm-org/LightGBM - Efficient Document Ranking with Learnable Late Interactions: https://arxiv.org/abs/2406.17968 - Passage Re-ranking with BERT: https://arxiv.org/abs/1901.04085 - A Thorough Comparison of Cross-Encoders and LLMs for Reranking SPLADE: https://arxiv.org/abs/2403.10407 - Context-1 Research Report: https://www.trychroma.com/research/context-1 Chapters 0:00 - Intro 2:37 - Example application signals and metrics 6:49 - Aggregating metrics 8:04 - Example algorithmic reranker in code — Chroma builds open-source search infrastructure for AI Fast, serverless, and scalable infrastructure supporting vector, full-text, regex, and metadata search. Built on object storage and trusted by millions of developers. Open-source Apache 2.0. https://trychroma.com/
Lexical Search in Chroma | Full Text Search, BM25 & SPLADE
4:41
Tutorial4:41
Lexical Search in Chroma | Full Text Search, BM25 & SPLADE
Chroma offers a variety of lexical search strategies. Full Text Search (FTS), BM25 and SPLADE are the core offerings. Each have their strengths and weaknesses and use cases. Docs on FTS: https://docs.trychroma.com/docs/querying-collections/full-text-search Docs on BM25: https://docs.trychroma.com/integrations/embedding-models/chroma-bm25#chroma-bm25 Docs on SPLADE: https://docs.trychroma.com/integrations/embedding-models/chroma-cloud-splade#chroma-cloud-splade 0:00 - Lexical search overview 0:30 - FTS 0:52 - BM25 1:23 - SPLADE 1:53 - Demo -- Chroma builds open-source data infrastructure for AI Fast, serverless, and scalable infrastructure supporting vector, full-text, regex, and metadata search. Built on object storage and trusted by millions of developers. Open-source Apache 2.0. https://trychroma.com/

Who uses it

Real use cases, no hype
Semantic Search
Use vector embeddings to find relevant documents by meaning, not just keywords.
Retrieval-Augmented Generation
Feed LLMs contextual data from your knowledge base for accurate, grounded answers.
Code Search
Index codebases with AST-aware chunking and power developer productivity tools.
Multi-Modal Retrieval
Search across images, audio, and text using unified embeddings.
Full-Text & Regex Search
Run fast keyword and regex queries without requiring an embedding model.
Agentic Search
Build agents that iteratively search and refine results to answer complex queries.
Metadata Filtering
Combine semantic search with structured filters for faceted retrieval.
What's great
  • Open-source Apache 2.0 with no vendor lock-in
  • Supports vector, full-text, regex, and metadata search
  • Easy to get started (pip install, 30 seconds)
  • Large active community (27k GitHub stars)
  • Serverless scaling with no manual ops
Technical strengths
  • Object-storage backed reduces cost by 10x
  • Automatic query-aware data tiering (hot/warm/cold)
Where it falls short
  • Write throughput limited to 30 MB/s per collection
  • No built-in distributed sharding (single node focus)
  • Query latency can spike on cold start
  • Limited to 5M records per collection
Technical limitations
  • Recall capped at 90-100% for some workloads
  • Heavy reliance on object storage introduces latency spikes
  • Some advanced features like multi-tenancy still maturing
For technical folks

The deep-dive specs

Architecture
Object-storage based with query-aware intelligent tiering
Write throughput (per collection)
30 MB/s (2000+ QPS)
Concurrent reads (per collection)
10 (200+ QPS)
Max collections per database
1M
Records per collection
5M
Recall
90-100%

FAQ

Includes technical Q&A
Yes, in most cases your credits do not expire. The $100 of included usage in the Team plan does not rollover.

Traffic Insights

Monthly Visits

July 2026

Growth Rate

0%

vs last month

Dominance

0%

in Search

Top Country

Top Countries

Growth Trend

Stable

Traffic has remained stable.

Honest pricing

No sneaky tiers, no “contact sales”

Here's who each plan is actually for — and where the hidden charges might hit.

Starter
Free

Individuals

  • $0 per month plus usage
  • 10 databases
  • 10 team members
  • Community Slack
  • $5 free credits
Team
Most picked
$250/mo

Teams

  • $250 per month plus usage
  • 100 databases
  • 30 team members
  • Slack support
  • SOC 2 Type II
  • Volume-based discounts
  • $100 included credits
Enterprise
Custom

Enterprise

  • Custom pricing
  • Unlimited databases
  • Unlimited team members
  • Dedicated support
  • Single tenant clusters
  • BYOC clusters
  • SLAs
Reviews

What the internet actually thinks

Data refreshed weekly. No paid placements.

Trustpilot
Not reviewed yet
G2
G2
Not reviewed yet
C
Capterra
Not reviewed yet
Product Hunt
Not reviewed yet
COMMUNITY COMMENTS

Leave a Comment

Help others make informed decisions. Your honest feedback shapes the community's understanding of this tool.

Tips for commenting

  • Be specific about features you used
  • Share real use cases and results
  • Mention both pros and cons
  • Keep it honest and constructive
Sign in to leave a comment

Your comment will be published after moderation

Filter by rating:
Sort by:

No comments yet

Be the first to share your experience with this tool.

Alternatives

Not quite the right fit?

Here's what else is out there.

Detail tags
Found via these searches
#VectorSearch#OpenSource#AI#MachineLearning#Search#Database#Embeddings