#weaviate
weaviate by @weaviate_io (⭐️ 16858)

Weaviate is an open-source vector database that stores both objects and vectors, allowing for the combination of vector search with structured filt...

#go
September 29, 2026 at 3:49 AM
Four vendors, one bad assumption: SSRF in MCP servers
There is a particular kind of bug that you only recognize the third time you see it. The first time, it looks like a mistake. The second time, it looks like a coincidence. By the third time you stop looking at the code and start looking at the people who wrote it, because the bug is no longer in the code. It is in an assumption everyone shared. Over the first nine months of 2026 I found the same assumption in MCP servers shipped by Google, Anthropic, Microsoft, and Weaviate. Four vendors, four codebases, four different languages and frameworks and review cultures. One bug shape. This is the story of that shape, how it got there, and what it took to get it out. ## What an MCP server actually is If you have not spent time with the Model Context Protocol, the one-sentence version is this: an MCP server is a program that accepts structured input from a language model and does something in the real world with it. Query a database. Fetch a web page. Drive a browser. Call an embedding API. That sentence hides the entire problem. "Accepts input from a language model" sounds like a closed loop. The model is your model. The server is your server. Who is the attacker? The answer is: whoever controls what the model reads. A prompt injected through a web page, a document, a support ticket, or a database row can steer the model, and the model will then steer the server. The model is not a trusted caller. It is a proxy for every untrusted input it has ever seen. Once you internalize that, every parameter an MCP server accepts becomes attacker-controlled by definition. The four vendors below had not internalized it. And the parameter they all forgot to distrust was a URL. ## Google: a redirect nobody checked Google's MCP Toolbox for Databases has an HTTP source type. You point it at a base URL, and tools built on that source make requests to paths under it. The code that builds the HTTP client lives in `internal/sources/http/http.go`. It created a standard Go `http.Client` with no `CheckRedirect` hook and no validation of where a request would actually land. That is two missing things, and they compound. Go's default client follows redirects on its own. Without a `CheckRedirect` policy, the toolbox had no opportunity to inspect each hop. Without target IP validation, even the first request could be pointed somewhere it should never go. Google assigned it CVE-2026-14540, published 2026-07-31 through Google's CNA. CVSS v4.0 8.0 High, CWE-918, Server-Side Request Forgery. Affected versions 0.3.0 through 1.4.0. The fix landed in googleapis/mcp-toolbox PR #3448, merged 2026-06-18 and released as v1.5.0 the same day, adding redirect and target validation. I was credited as the finder. I wrote a full article on this one. The short version for this piece: the assumption was that the base URL configured by the operator was the whole trust decision, and everything after that was safe by inheritance. Redirects broke that inheritance. ## Anthropic and Microsoft: the fetch that skips its own guard Anthropic's reference `mcp-server-fetch` and Microsoft's `playwright-mcp` do different jobs. One retrieves a URL and returns its content. The other drives a real browser. Both take a URL from the model. Neither had an allowlist. Neither blocked internal IP ranges. Neither filtered addresses that resolve to link-local, loopback, or private space. A model that has been steered can ask either one to fetch an internal admin panel or a metadata endpoint, and the server will comply. The `mcp-server-fetch` case had a sharper edge. The server does contain a safeguard, a function called `check_may_autonomously_fetch_url()` that is meant to gate what the server will retrieve on the model's behalf. But the `get_prompt` handler calls `fetch_url()` directly and never invokes the check. The guard exists. There is a code path around it. This is a shape I have come to expect in MCP servers. A security control gets added to the primary tool-call path, and the secondary paths (prompts, resources, completions) are written by someone else or on a different day and simply do not route through it. I published both issues together on the Full Disclosure mailing list on 2026-05-25. CVSS 3.1 7.5. Neither issue was a secret when I posted, both were already visible in public GitHub threads. What the disclosure did was consolidate them, assign severity, and put them in a place where operators would actually see them. As of the disclosure, I cannot confirm that either vendor has shipped a fix, and I am not going to claim otherwise here. ## Weaviate: the field that dodged two hardening passes This one is my favorite, because Weaviate had already fixed this bug. Twice. Just not for the field I found. Weaviate is a vector database with pluggable modules that call out to embedding and generation APIs. The Google-backed modules (`text2vec-google`, `multi2vec-google`, `generative-google`) send requests to Google's API with the operator's Google API key as a bearer credential. Under `USE_GOOGLE_AUTH=true`, they instead send a live GCP OAuth token scoped to cloud-platform. Most Weaviate modules let you override the upstream host through a field called `baseURL`. Weaviate had hardened that field in two passes: PR #10878, merged 2026-03-27, and PR #11683, merged 2026-06-18, covering 21 URL builders. The Google modules do not call their field `baseURL`. They call it `apiEndpoint`. Both hardening passes were structurally scoped to `baseURL`, so `apiEndpoint` sailed through both of them untouched. The consequence: a user who could set `apiEndpoint` could point the module at a host they controlled and receive the operator's Google API key, or the operator's GCP OAuth token, in the request. There were two ways to set it: through the class schema config, which requires schema-write access, and through a GraphQL query-time parameter, reachable with ordinary read access. You did not need to be an administrator. You needed to be able to run a query. I reported it through HackerOne. Weaviate Security confirmed it. The fix is weaviate/weaviate PR #12961, merged 2026-09-07 into stable/v1.37. Weaviate credited me for the report. ## The assumption Line up the four cases and the shared assumption is obvious in hindsight. Google assumed the operator-configured base URL settled the trust question, and that redirects inherited that trust. Anthropic and Microsoft assumed the URL a model asks for is a URL the model should get, and in one case wrote a check but did not wire it to every path. Weaviate assumed that hardening `baseURL` meant hardening "the field that controls the upstream host," when one module family had spelled that field differently. In every case, a URL crossed a trust boundary and nobody was standing at the boundary. That is what SSRF is. What makes MCP different is that the trust boundary has moved. In a classic web app, the attacker types the URL into a form. In an MCP server, the attacker plants text somewhere a model will read it, and the model types the URL. The server sees a request from its own trusted model and does not think to ask where the idea came from. ## How I found the pattern I did not find these by reading four codebases end to end. I found them because I had stopped being able to read MCP servers by hand and had written a tool to do the first pass for me. `mcp-safeguard` is an open-source static analysis scanner for MCP servers. It is on PyPI, MIT licensed, and it runs around 150 rules across seven categories: prompt injection, credential leaks, endpoint exposure, tool poisoning, SSRF, OAuth scope, and source audit. When the same rule fires on Google's Go code and Anthropic's Python and Weaviate's module layer, you stop treating each hit as a one-off. I generalized the findings into an IETF Internet-Draft, `draft-mohiuddin-mcp-security-considerations-00`, which lays out six vulnerability classes and names the underlying move "Protocol Pivoting": an attacker enters through the model-facing protocol and pivots into whatever the server can reach behind it. ## Where this stands Four confirmed findings. Google fixed and issued a CVE. Weaviate fixed and credited. Anthropic and Microsoft disclosed publicly, with fix status unconfirmed as of that disclosure. And the tool that surfaced the pattern is public and free. There is a fifth report, still working through a vendor's disclosure process. It will be added here once it is public. ## Timeline * 2026-03-27: Weaviate PR #10878 (baseURL validation, opt-in) merged * 2026-05-25: Anthropic mcp-server-fetch and Microsoft playwright-mcp SSRF disclosed on Full Disclosure * June 2026: IETF Internet-Draft draft-mohiuddin-mcp-security-considerations-00 published * 2026-06-18: Google PR #3448 (SSRF guard) merged, v1.5.0 released * 2026-06-18: Weaviate PR #11683 (X-*-BaseURL header validation) merged * 2026-07-31: CVE-2026-14540 published by Google's CNA * 2026-09-07: Weaviate PR #12961 (Google module apiEndpoint restriction) merged into stable/v1.37 ## References * CVE-2026-14540 * Google fix, googleapis/mcp-toolbox PR #3448 * Full Disclosure post, 2026-05-25 * modelcontextprotocol/servers #4116, #4143, #4205 * microsoft/playwright-mcp #1626 * Weaviate PR #10878 * Weaviate PR #11683 * Weaviate fix, PR #12961 * mcp-safeguard * IETF Internet-Draft _Syed Anas Mohiuddin is an AI security researcher focused on Model Context Protocol security and the founder of Cognivators. Portfolio: https://syedanas01.github.io/ · GitHub: https://github.com/SyedAnas01 · ORCID: https://orcid.org/0009-0005-3736-6430_
dev.to
September 29, 2026 at 3:49 AM
Weaviate outlines how late-interaction multi-vector models and page-level embeddings let RAG pipelines ingest and query charts, tables, and diagrams in PDFs.
September 27, 2026 at 3:28 PM
Weaviate 1.39 reshapes vector search efficiency.

Introducing 4-bit Rotational Quantization (RQ4/RQ4c) and SIMD FWHT, Weaviate 1.39 slashes encoder latency and index size, while sustaining RAG recall. Key gains: 32-37% faster imports, ~45% less heap use. Critical for…

Read more on Kimbodo:
Retrieval, RAG & Search — September 17, 2026
What Happened Weaviate 1.39 introduced 4-bit Rotational Quantization (RQ4, plus an uncentered variant RQ4c) and a SIMD Fast Walsh–Hadamard Transform implementation that significantly reduces encoder latency, index size and import…
kimbodo.com
September 24, 2026 at 8:30 PM
Vector Database Comparison: 7 Self-Hosted Tools—What Actually Differs

Compare Chroma, Qdrant, Weaviate, Milvus, Vespa, Vald, LanceDB: RAM, license, offline capability, maturity. Pick the right one.

https://forgedgoods.org/g/vector-database-comparison-self-hosted-tools.html
September 24, 2026 at 10:00 AM
Retrieval-Augmented Generation (RAG) evolves with production-ready tools.

RAG now features standard production patterns like embedding pipelines and hybrid BM25+vector retrieval. Choices include managed and open-source vector databases like Pinecone and Weaviate. Enhanced…

Read more on Kimbodo:
How to Build Reliable Retrieval‑Augmented Generation: Vector DB choices, architectures and production best practices
What Happened Retrieval‑Augmented Generation (RAG) has moved from prototypes to production patterns: embedding pipelines, ANN vector stores, hybrid BM25+vector retrieval, and reranking are now standard. Frameworks and orchestration layers (LlamaIndex,…
kimbodo.com
September 22, 2026 at 5:15 PM
Let’s break down the core differences every AI engineer must know:

🔹 Traditional Vector RAG
• How it works: Document -> Chunking -> Vector Embeddings -> Vector DB (Pinecone/Weaviate) -> Similarity Search -> Top-K Chunks -> LLM Output.
September 19, 2026 at 3:45 AM
Elasticsearch and Weaviate shift RAG & search choices.

Elasticsearch now offers a serverless vector DB with hybrid search and BBQ compression, enabling storage of billions of vectors. Weaviate's HFresh minimizes heap usage with disk-backed vector indexing. These updates…

Read more on Kimbodo:
Retrieval, RAG & Search — September 9, 2026
What Happened Three technology developments change practical choices for retrieval‑augmented generation (RAG) and semantic search: Elasticsearch launched a serverless Elasticsearch Vector Database with vector‑first index modes, built‑in hybrid search, managed…
kimbodo.com
September 16, 2026 at 4:26 PM
Ten open-source repos for RAG retrieval: chroma, qdrant, weaviate, milvus, lancedb, llama_index, datahub, WeKnora, markitdown, and OpenViking.
September 13, 2026 at 3:35 PM
30. AnythingLLM

31. Pinecone
32. Weaviate
33. Qdrant
34. Chroma
35. Supabase
36. MongoDB Atlas
37. PostgreSQL + pgvector
38. Elasticsearch
39. Zapier
40. Make

41. n8n
42. Flowise
43. Dify
44. Streamlit
45. Gradio
46. Poe
47. Midjourney
48. Stable Diffusion
49. ElevenLabs
September 12, 2026 at 11:30 AM
RAG matures with orchestration libraries and vector stores.

Retrieval-augmented generation (RAG) has evolved: LlamaIndex, LangChain, and Haystack now orchestrate chunking, embeddings, and neural reranking. Platforms like Pinecone, Qdrant, and Weaviate enhance performance…

Read more on Kimbodo:
Retrieval, RAG & Search — September 2, 2026
What Happened Retrieval-augmented generation (RAG) is now a mature pattern: application logic orchestrates chunking, embeddings, ANN search, metadata filtering and neural reranking to provide high-precision grounding for LLMs. The ecosystem…
kimbodo.com
September 10, 2026 at 5:10 PM
Dify 1.17.1

Self-hosted deployments using the bundled Weaviate must complete a manual, staged upgrade before starting 1.17.1. The bundled Weaviate server moves from 1.27.0 to 1.39.2 — 12 minor versions — and skipping minors is unsupported. Pulling and restarting can silently and permanently break…
Dify 1.17.1
Self-hosted deployments using the bundled Weaviate must complete a manual, staged upgrade before starting 1.17.1. The bundled Weaviate server moves from 1.27.0 to 1.39.2 — 12 minor versions — and skipping minors is unsupported. Pulling and restarting can silently and permanently break vector…
whatsnew.fyi
September 10, 2026 at 4:02 PM
weaviate by @weaviate_io (⭐️ 16793)

Weaviate is an open-source vector database that stores both objects and vectors, allowing for the combination of vector search with structured filt...

#go
September 8, 2026 at 4:11 AM
Vector Database Selection Guide: Pinecone, Weaviate, and Chroma Compared

Discover how to pick the right vector database for your AI workloads. A deep dive into Pinecone, Weaviate, and Chroma with feature…

https://ai-blog-seven-wine.vercel.app/en/posts/2026-09-08-am-pwf3o

#vector-database #AI #RAG
September 8, 2026 at 2:35 AM
Weaviate leads today's durable movers. Heat score: 65 ▲16 in 24h, +45 over 7 days. The vector database framework is holding momentum, not just spiking. #AIFrameworks

https://hookflow.ai/tools/weaviate
HookFlow.ai — Best Trending AI Tools, Ranked by Live Heat Scores
Discover the best trending AI tools in real-time. HookFlow ranks 711+ AI tools by live heat scores from Reddit, GitHub, and Product Hunt — updated daily.
hookflow.ai
September 7, 2026 at 12:16 PM
Weaviate hits heat score 49 on HookFlow, up 19 pts in 24h and 33 pts over 7 days. The vector DB is sustaining momentum — not a spike. #AIFrameworks ▲

https://hookflow.ai/tools/weaviate
HookFlow.ai — Best Trending AI Tools, Ranked by Live Heat Scores
Discover the best trending AI tools in real-time. HookFlow ranks 708+ AI tools by live heat scores from Reddit, GitHub, and Product Hunt — updated daily.
hookflow.ai
September 6, 2026 at 1:03 PM
ベクトルデータベース選定ガイド ― Pinecone・Weaviate・Chroma 徹底比較と活用事例

Pinecone、Weaviate、Chroma の特徴・スケーラビリティ・運用コストを徹底比較。実際の導入事例と選定フローで、AIシステムに最適なベクトルDBを見つける方法を解説します。

https://ai-blog-seven-wine.vercel.app/ja/posts/2026-09-06-am-npq1r

#ベクトルデータベース #AI #RAG
September 6, 2026 at 2:25 AM