展示 HN:查找陈旧、孤立且已删除但可恢复的 RAG 向量
Show HN: Find stale, orphaned, deleted-but-retrievable RAG vectors

原始链接: https://github.com/rimironenko/rag-staleness-check

**rag-staleness-check** 是一款开源的只读诊断工具,专为审计 RAG 向量索引(pgvector、Qdrant 和 Chroma)的数据合规性问题而设计。通过将已索引的数据块与源清单进行比对,它可以识别以下问题: * **陈旧数据 (Staleness):** 源文档已被修改但索引尚未更新的数据块。 * **孤儿数据 (Orphans):** 源文档已不存在但仍留在索引中的数据块。 * **重复数据 (Duplicates):** 通过余弦相似度或精确内容哈希识别出的近乎相同的数据块。 * **可检索性 (Retrievability):** 探测“功能性泄漏”,确认已删除的 ID 是否在物理上仍可访问或可被索引检索。 该工具优先考虑安全性,通过结构化语法分析和会话级配置(例如 Postgres 的 `default_transaction_read_only`)强制执行只读操作。它在本地运行,默认不进行任何遥测,也不会执行任何写入、删除或更新操作。 用户需提供一份当前文档的清单,并可选择提供一份已删除 ID 的列表。该工具会将结果评分发送至标准输出 (stdout),并生成详细的 JSON 报告。虽然它提供针对特定引擎的诊断,但不具备 premium 级私有 RAGproof 审计套件所特有的多引擎编排或账本验证精度。可通过 `pip` 安装,并针对特定向量数据库提供模块化扩展。

```Hacker News新帖 | 往期 | 评论 | 提问 | 展示 | 招聘 | 提交登录Show HN: 查找陈旧、孤立且可恢复的已删除 RAG 向量 (github.com/rimironenko)由 rimironenko 在 1 小时前发布,3 点 | 隐藏 | 往期 | 收藏 | 讨论 帮助 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系方式 搜索: ```
相关文章

原文

Read-only staleness / orphan / duplicate / retrievability checks for a single pgvector, Qdrant or Chroma index - the open-source half of the RAGproof "decayed RAG index" teardown (full writeup + multi-engine ledger-verified methodology).

This tool runs against your own already-indexed vector database and tells you:

  • staleness - indexed chunks whose source document has since changed (needs a manifest with last_modified per document and the engine itself storing a per-row last-modified value)
  • orphans - indexed chunks whose source document no longer exists in your manifest at all
  • duplicates - near-identical chunks (cosine similarity ≥ threshold, default 0.98), plus an exact-hash pass if you store a per-chunk content hash
  • retrievable-after-delete - a basic probe: for ids you believe are deleted, is the vector still physically retrievable by id (storage-layer persistence), and does it still surface in top-k search results (functional leak)?

It's read-only - it never calls a write/delete/upsert method on your engine. See Read-only guarantee below for how that's enforced, and its limits.

What this tool does NOT do

This is the single-engine, ledger-free, self-serve slice. It doesn't include:

  • multi-engine orchestration (run it once per engine yourself)
  • the GDPR Article-17 reporting pack (signed evidence, verifiable-erasure report)
  • engine-specific deep checks (pgvector dead-tuple/VACUUM detail, Qdrant optimizer-threshold detail, Chroma on-disk HNSW growth)
  • a ledger-based precision/recall harness - there's no ground truth here, this tool reports what it finds, not how accurate the finding is

Those live in the private, paid RAGproof audit, which cross-validates every finding against a git-history-derived ground-truth ledger across all three engines simultaneously.

pip install rag-staleness-check              # pgvector only
pip install "rag-staleness-check[qdrant]"     # + Qdrant
pip install "rag-staleness-check[chroma]"     # + Chroma
pip install "rag-staleness-check[all]"        # everything

Or, without installing:

pipx run rag-staleness-check --engine pgvector --dsn "$DSN" --source ./docs_manifest.json

Requires Python 3.10+.

rag-staleness-check \
  --engine pgvector \
  --dsn "postgresql://readonly_user:pw@localhost:5432/mydb" \
  --pg-table chunks \
  --pg-doc-id-column doc_id \
  --pg-last-modified-column last_modified \
  --pg-content-hash-column content_sha256 \
  --source ./docs_manifest.json \
  --deleted-ids ./deleted_ids.json \
  --out findings.json

This prints a scorecard to stdout and writes the full result to --out (default findings.json). If a check is missing a required input - no --source, no per-row metadata configured, no --deleted-ids - it reports "skipped": true with a reason instead of silently showing a misleadingly clean 0%.

Engine --dsn format Example
pgvector Postgres DSN postgresql://user:pw@localhost:5432/db
qdrant Base URL http://localhost:6333
chroma host:port localhost:8000

Your table/column or payload/metadata field names are your own - this tool doesn't assume anything about your schema beyond what you tell it via flags:

pgvector (--pg-*): --pg-table (required), --pg-id-column (default id), --pg-vector-column (default embedding), --pg-doc-id-column, --pg-last-modified-column, --pg-content-hash-column. The last three are optional, and leaving them out is exactly what makes staleness / orphans / the duplicates check's exact-hash pass report skipped / zero-coverage.

Qdrant / Chroma (--collection required, plus payload/metadata field names): --doc-id-field (default doc_id), --last-modified-field (default last_modified), --content-hash-field (default content_sha256).

A JSON file describing what documents should exist - this is the join key for staleness and orphans. There's no ledger in this tool; the manifest is the only source of truth about "what's current."

{
  "manifest_version": 1,
  "documents": [
    {
      "doc_id": "handbook/engineering/on-call.md",
      "last_modified": "2026-06-01T00:00:00Z",
      "content_sha256": "optional, doc-level, informational only"
    }
  ]
}

doc_id has to match whatever value your own ingestion pipeline wrote into each indexed chunk's doc-id column/field (see schema mapping above). last_modified is only needed if you want the staleness check to run. content_sha256 here is doc-level and purely informational - it isn't the input to duplicate detection (see below). Full example at examples/docs_manifest.json.

Deleted-ids probe (--deleted-ids)

A JSON array of ids you believe are deleted:

["chunk-abc123", "chunk-def456"]

For each one, this checks whether the vector is still retrievable by id (storage-layer persistence) and, if so, whether it still shows up in a top-k self-query (functional leak). The distinction matters: per EDPB Guidelines 05/2019, erasure has to be verifiable and irreversible - suppressing a record from search results isn't enough on its own if the underlying vector is still there. findings.json's retrievable.framing field spells this out every run.

This tool never calls a write/delete/upsert method on your engine, and that's enforced two ways:

  1. Structurally, via a static test - tests/test_no_write_methods.py walks the actual syntax tree of every file under src/rag_staleness_check/ and fails the build if any write/delete/ upsert-shaped method gets called, or any SQL write statement is passed to execute()/executemany(). Not a substring grep - that would false-positive on this very README and on docstrings - it parses real syntax.
  2. At runtime, for pgvector - every connection opens with SET default_transaction_read_only = on, and rag_staleness_check.readonly.assert_pg_read_only() re-checks SHOW default_transaction_read_only before any query runs, refusing to proceed if it isn't 'on'. That's defense-in-depth against a connection pooler (e.g. PgBouncer) silently swallowing the session-level SET.

Qdrant and Chroma don't give a client-exposed way to check "is this session/API key read-only" - that's deployment-side RBAC, not something either client library reports. For those two engines, enforcement is the static test above plus your own deployment-side scoping: connect with a read-only-scoped API key/token if your engine supports one. There's no runtime assertion for Qdrant/Chroma equivalent to pgvector's - flagging that here rather than implying parity where there isn't any.

For extra assurance on pgvector, connect through a locked-down role instead of relying on this tool alone:

CREATE ROLE rag_staleness_check_readonly NOSUPERUSER LOGIN PASSWORD '...';
GRANT CONNECT ON DATABASE mydb TO rag_staleness_check_readonly;
GRANT USAGE ON SCHEMA public TO rag_staleness_check_readonly;
GRANT SELECT ON chunks TO rag_staleness_check_readonly;

No default telemetry. Nothing is collected or sent automatically. --share-anonymous-scorecard is an explicit opt-in that prints the exact anonymized payload (aggregate counts + engine type only - never your dsn, hostnames, doc_ids, or chunk_ids) that would be shared. There's no backend configured yet, so nothing actually goes over the network - the flag is reserved for a future opt-in submission endpoint. DO_NOT_TRACK in the environment, if set, forces this off even if the flag is passed.

Separately, this tool disables chromadb's own client-side telemetry (anonymized_telemetry=False) when connecting to Chroma. The chromadb package ships its own default-on PostHog-based telemetry, independent of anything above - leaving it enabled would quietly break the "no default telemetry" promise for --engine chroma users even though this tool's own code never phones home.

One cosmetic wrinkle: on chromadb==0.6.3, the setting takes effect (verified via client._system.settings.anonymized_telemetry is False), but a startup event (ClientStartEvent) still tries to fire through a code path that doesn't consistently honor it, and throws capture() takes 1 positional argument but 3 were given while constructing the call - before any request is made. You may see this line printed; it doesn't mean anything was sent.

  • Duplicate detection's exact-hash pass needs a per-chunk content hash stored in your engine (--pg-content-hash-column / --content-hash-field). Without one, it falls back to cosine-ANN only, and a warning field in the duplicates check says so rather than staying silent about it.
  • Chroma's snapshot_stats() only reports a live count, intentionally
    • no on-disk HNSW-directory-size measurement. (The private RAGproof audit's dev-environment build of this used a docker exec du trick against a known-local container; that has no equivalent against a real third-party deployment, so it isn't shipped here.)
  • qdrant-client's own compatibility check may warn if your installed client's minor version differs from your server's by more than 1 (e.g. "Qdrant client version 1.18.0 is incompatible with server version 1.15.5"). Non-fatal - the check still runs - but if it's noisy, pin qdrant-client to match your server's minor version yourself.
  • This tool reports counts and pairs, not precision/recall. There's no ground truth available client-side to score itself against - that's what the private, ledger-based audit is for.
  • Duplicate detection needs a retrievable vector per candidate - a chunk with no vector returned by get_vector is silently excluded from the cosine-ANN pass (nothing to query with), not counted as "not a duplicate."
  • qdrant-client and chromadb are optional extras (pip install rag-staleness-check[qdrant] / [chroma] / [all]) so a pgvector-only user doesn't have to pull in Chroma's heavier transitive dependencies.
  • Neither is pinned to an exact version. This tool connects to a third party's already-running deployment, whose server version is out of its control, so hard-pinning - unlike a project that ships and controls its own Docker images - would just break users on anything else.

Apache-2.0. See LICENSE. Independent of any other RAGproof project's license.

Issues and PRs welcome. Not yet implemented: a --method minhash surface-text-based duplicate detection mode (via datasketch) as an alternative/complement to the cosine-embedding approach above - PRs welcome.

联系我们 contact @ memedata.com