Vol. XVI · No. 261Friday 18 September 2026World Edition
TheNewsRupt coat of arms crest

The NewsRupt

ai recommendation index

AI Recommendation Index — Infrastructure, Q3 2026

Q3 2026v1.0Collected 2026-09-182026-09-18Published 18 Sep 2026

We put 12 buying questions about vector databases, agent frameworks and inference providers to three models, three times each — 108 sampled answers — and recorded every vendor named. Two findings stand out: recommendation is heavily concentrated at the top of each category, and the first name an assistant gives is unstable. In 36% of model-prompt cells, the same model asked the identical question three times led with a different vendor each time.

What we found
  1. Concentration is extreme in agent frameworks

    LangGraph was named in 100% of the 36 agent-framework answers and led 88.9% of them — the highest index score in the study at 95.6. No other category had a name approaching that dominance.

  2. Mention and first-place are different races

    Weaviate was mentioned in 80.6% of vector-database answers — more often than Pinecone at 69.4% — but led only 8.3% of them against Pinecone's 50%. Being on the list and being the recommendation are separate outcomes.

  3. The lead name is unstable

    Across 36 model-prompt cells sampled three times each, 13 (36%) produced more than one distinct first-named vendor. Instability was 17% for gemini-3.8-flash, 42% for gemini-3.1-flash-lite and 50% for gpt-5.4-mini.

  4. Models disagree about the second tier

    Replicate was named in 75% of gpt-5.4-mini inference answers and 0% by both Gemini models. Semantic Kernel appeared in 83.3% of gpt-5.4-mini agent answers and 8.3% of gemini-3.1-flash-lite ones. Single-model visibility audits will mislead.

Agent frameworksIndex score 0–100
VendorIndexMentionedNamed firstgpt-5.4-minigemini-3.8-flashgemini-3.1-flash-lite
LangGraph95.6100%88.9%100%100%100%
LangChain5083.3%0%50%100%100%
Semantic Kernel32.847.2%11.1%83.3%50%8.3%
AutoGen2541.7%0%50%25%50%
CrewAI21.736.1%0%25%33.3%50%
LlamaIndex11.719.4%0%50%8.3%0%
Haystack6.711.1%0%16.7%0%16.7%
Temporal6.711.1%0%16.7%16.7%0%
Agno1.72.8%0%0%8.3%0%
Inference providersIndex score 0–100
VendorIndexMentionedNamed firstgpt-5.4-minigemini-3.8-flashgemini-3.1-flash-lite
Together AI71.194.4%36.1%83.3%100%100%
Fireworks AI42.869.4%2.8%83.3%83.3%41.7%
vLLM31.144.4%11.1%16.7%58.3%58.3%
Groq28.938.9%13.9%41.7%25%50%
Baseten21.127.8%11.1%33.3%41.7%8.3%
Anyscale19.430.6%2.8%16.7%8.3%66.7%
RunPod18.330.6%0%8.3%41.7%41.7%
Modal16.727.8%0%33.3%8.3%41.7%
Replicate1525%0%75%0%0%
Deepinfra13.316.7%8.3%16.7%33.3%0%
Vertex AI12.216.7%5.6%41.7%8.3%0%
Lambda Labs11.719.4%0%0%33.3%25%
Hugging Face8.313.9%0%25%16.7%0%
Amazon Bedrock6.18.3%2.8%8.3%16.7%0%
Azure AI6.18.3%2.8%8.3%16.7%0%
Cerebras6.18.3%2.8%0%25%0%
Vector databasesIndex score 0–100
VendorIndexMentionedNamed firstgpt-5.4-minigemini-3.8-flashgemini-3.1-flash-lite
Pinecone61.769.4%50%66.7%66.7%75%
Weaviate51.780.6%8.3%91.7%58.3%91.7%
Qdrant48.369.4%16.7%66.7%100%41.7%
Milvus36.750%16.7%50%50%50%
Elasticsearch2027.8%8.3%33.3%25%25%
OpenSearch16.727.8%0%50%33.3%0%
pgvector16.727.8%0%8.3%58.3%16.7%
Zilliz13.322.2%0%16.7%33.3%16.7%
MongoDB Atlas8.313.9%0%25%0%16.7%
Vespa6.711.1%0%33.3%0%0%
Chroma1.72.8%0%0%0%8.3%
Method
Models
google/gemini-3.1-flash-lite, google/gemini-3.8-flash, openai/gpt-5.4-mini
Repetitions per prompt
3
Total sampled answers
108
Sampling
Each prompt was sent independently to each model three times with no conversation history, no system prompt beyond the question, and no retrieval tools.
Extraction
Vendor names were extracted from each answer by case-insensitive string matching against a fixed candidate list per category, plus recorded alias spellings. The first candidate to appear in the text was recorded as the first-named vendor.
Scoring
Index score = 60% x mention rate + 40% x first-named rate, expressed on a 0-100 scale. Mention rate is the share of answers in which the vendor appeared; first-named rate is the share in which it appeared first.
The frozen prompt set
Agent frameworks
  • Which framework should I use to build a production AI agent that calls internal APIs?
  • What is the best Python library for orchestrating multi-step LLM agents?
  • I want to build an agent with tool calling and human approval steps. What should I use?
  • Which agent framework is the most stable choice for an enterprise team in 2026?
Vector databases
  • Which vector database should I use for a production RAG system with about 50 million embeddings?
  • What vector database do you recommend for a small engineering team that wants managed hosting?
  • I need hybrid keyword and vector search in one system. What should I use?
  • What are the best options for self-hosted vector search on Kubernetes?
Inference providers
  • Which inference provider should I use to serve an open-weight model cheaply at scale?
  • What is the best API provider for low-latency open-source model inference?
  • Where should I host a fine-tuned Llama-class model for production traffic?
  • Which provider do you recommend for high-throughput batch LLM inference?
Limitations
  • Three models are not the market. The index covers two Google models and one OpenAI model served through one gateway; assistants with live web retrieval, memory or a system prompt will behave differently.
  • Vendor extraction is string matching against a fixed candidate list, so a vendor named only by an unlisted alias is undercounted, and a vendor named only to be dismissed still counts as a mention.
  • Twelve prompts cannot represent all buying language for these categories. Phrasing changes results, which is part of the finding rather than a controlled variable.
  • Three repetitions per cell is enough to demonstrate instability but not to estimate its rate precisely; treat the 36% figure as a floor with wide error.
  • This measures what assistants say, not product quality. Nothing here is a recommendation.
Suggested citation

The NewsRupt, “AI Recommendation Index — Infrastructure, Q3 2026” (v1.0), published 18 Sep 2026. Data licensed CC BY 4.0.