AI Recommendation Index — Infrastructure, Q3 2026
We put 12 buying questions about vector databases, agent frameworks and inference providers to three models, three times each — 108 sampled answers — and recorded every vendor named. Two findings stand out: recommendation is heavily concentrated at the top of each category, and the first name an assistant gives is unstable. In 36% of model-prompt cells, the same model asked the identical question three times led with a different vendor each time.
Concentration is extreme in agent frameworks
LangGraph was named in 100% of the 36 agent-framework answers and led 88.9% of them — the highest index score in the study at 95.6. No other category had a name approaching that dominance.
Mention and first-place are different races
Weaviate was mentioned in 80.6% of vector-database answers — more often than Pinecone at 69.4% — but led only 8.3% of them against Pinecone's 50%. Being on the list and being the recommendation are separate outcomes.
The lead name is unstable
Across 36 model-prompt cells sampled three times each, 13 (36%) produced more than one distinct first-named vendor. Instability was 17% for gemini-3.8-flash, 42% for gemini-3.1-flash-lite and 50% for gpt-5.4-mini.
Models disagree about the second tier
Replicate was named in 75% of gpt-5.4-mini inference answers and 0% by both Gemini models. Semantic Kernel appeared in 83.3% of gpt-5.4-mini agent answers and 8.3% of gemini-3.1-flash-lite ones. Single-model visibility audits will mislead.
| Vendor | Index | Mentioned | Named first | gpt-5.4-mini | gemini-3.8-flash | gemini-3.1-flash-lite |
|---|---|---|---|---|---|---|
| LangGraph | 95.6 | 100% | 88.9% | 100% | 100% | 100% |
| LangChain | 50 | 83.3% | 0% | 50% | 100% | 100% |
| Semantic Kernel | 32.8 | 47.2% | 11.1% | 83.3% | 50% | 8.3% |
| AutoGen | 25 | 41.7% | 0% | 50% | 25% | 50% |
| CrewAI | 21.7 | 36.1% | 0% | 25% | 33.3% | 50% |
| LlamaIndex | 11.7 | 19.4% | 0% | 50% | 8.3% | 0% |
| Haystack | 6.7 | 11.1% | 0% | 16.7% | 0% | 16.7% |
| Temporal | 6.7 | 11.1% | 0% | 16.7% | 16.7% | 0% |
| Agno | 1.7 | 2.8% | 0% | 0% | 8.3% | 0% |
| Vendor | Index | Mentioned | Named first | gpt-5.4-mini | gemini-3.8-flash | gemini-3.1-flash-lite |
|---|---|---|---|---|---|---|
| Together AI | 71.1 | 94.4% | 36.1% | 83.3% | 100% | 100% |
| Fireworks AI | 42.8 | 69.4% | 2.8% | 83.3% | 83.3% | 41.7% |
| vLLM | 31.1 | 44.4% | 11.1% | 16.7% | 58.3% | 58.3% |
| Groq | 28.9 | 38.9% | 13.9% | 41.7% | 25% | 50% |
| Baseten | 21.1 | 27.8% | 11.1% | 33.3% | 41.7% | 8.3% |
| Anyscale | 19.4 | 30.6% | 2.8% | 16.7% | 8.3% | 66.7% |
| RunPod | 18.3 | 30.6% | 0% | 8.3% | 41.7% | 41.7% |
| Modal | 16.7 | 27.8% | 0% | 33.3% | 8.3% | 41.7% |
| Replicate | 15 | 25% | 0% | 75% | 0% | 0% |
| Deepinfra | 13.3 | 16.7% | 8.3% | 16.7% | 33.3% | 0% |
| Vertex AI | 12.2 | 16.7% | 5.6% | 41.7% | 8.3% | 0% |
| Lambda Labs | 11.7 | 19.4% | 0% | 0% | 33.3% | 25% |
| Hugging Face | 8.3 | 13.9% | 0% | 25% | 16.7% | 0% |
| Amazon Bedrock | 6.1 | 8.3% | 2.8% | 8.3% | 16.7% | 0% |
| Azure AI | 6.1 | 8.3% | 2.8% | 8.3% | 16.7% | 0% |
| Cerebras | 6.1 | 8.3% | 2.8% | 0% | 25% | 0% |
| Vendor | Index | Mentioned | Named first | gpt-5.4-mini | gemini-3.8-flash | gemini-3.1-flash-lite |
|---|---|---|---|---|---|---|
| Pinecone | 61.7 | 69.4% | 50% | 66.7% | 66.7% | 75% |
| Weaviate | 51.7 | 80.6% | 8.3% | 91.7% | 58.3% | 91.7% |
| Qdrant | 48.3 | 69.4% | 16.7% | 66.7% | 100% | 41.7% |
| Milvus | 36.7 | 50% | 16.7% | 50% | 50% | 50% |
| Elasticsearch | 20 | 27.8% | 8.3% | 33.3% | 25% | 25% |
| OpenSearch | 16.7 | 27.8% | 0% | 50% | 33.3% | 0% |
| pgvector | 16.7 | 27.8% | 0% | 8.3% | 58.3% | 16.7% |
| Zilliz | 13.3 | 22.2% | 0% | 16.7% | 33.3% | 16.7% |
| MongoDB Atlas | 8.3 | 13.9% | 0% | 25% | 0% | 16.7% |
| Vespa | 6.7 | 11.1% | 0% | 33.3% | 0% | 0% |
| Chroma | 1.7 | 2.8% | 0% | 0% | 0% | 8.3% |
- Models
- google/gemini-3.1-flash-lite, google/gemini-3.8-flash, openai/gpt-5.4-mini
- Repetitions per prompt
- 3
- Total sampled answers
- 108
- Sampling
- Each prompt was sent independently to each model three times with no conversation history, no system prompt beyond the question, and no retrieval tools.
- Extraction
- Vendor names were extracted from each answer by case-insensitive string matching against a fixed candidate list per category, plus recorded alias spellings. The first candidate to appear in the text was recorded as the first-named vendor.
- Scoring
- Index score = 60% x mention rate + 40% x first-named rate, expressed on a 0-100 scale. Mention rate is the share of answers in which the vendor appeared; first-named rate is the share in which it appeared first.
- “Which framework should I use to build a production AI agent that calls internal APIs?”
- “What is the best Python library for orchestrating multi-step LLM agents?”
- “I want to build an agent with tool calling and human approval steps. What should I use?”
- “Which agent framework is the most stable choice for an enterprise team in 2026?”
- “Which vector database should I use for a production RAG system with about 50 million embeddings?”
- “What vector database do you recommend for a small engineering team that wants managed hosting?”
- “I need hybrid keyword and vector search in one system. What should I use?”
- “What are the best options for self-hosted vector search on Kubernetes?”
- “Which inference provider should I use to serve an open-weight model cheaply at scale?”
- “What is the best API provider for low-latency open-source model inference?”
- “Where should I host a fine-tuned Llama-class model for production traffic?”
- “Which provider do you recommend for high-throughput batch LLM inference?”
- Three models are not the market. The index covers two Google models and one OpenAI model served through one gateway; assistants with live web retrieval, memory or a system prompt will behave differently.
- Vendor extraction is string matching against a fixed candidate list, so a vendor named only by an unlisted alias is undercounted, and a vendor named only to be dismissed still counts as a mention.
- Twelve prompts cannot represent all buying language for these categories. Phrasing changes results, which is part of the finding rather than a controlled variable.
- Three repetitions per cell is enough to demonstrate instability but not to estimate its rate precisely; treat the 36% figure as a floor with wide error.
- This measures what assistants say, not product quality. Nothing here is a recommendation.
The NewsRupt, “AI Recommendation Index — Infrastructure, Q3 2026” (v1.0), published 18 Sep 2026. Data licensed CC BY 4.0.
