Verifiable provenance
Every record carries its source URL, fetch timestamp and SHA-256. Every release ships a Merkle manifest signed via sigstore — check any record independently.
source_url · fetched_at · sha256
$ ragdata verify us-federal-procurement@2026.08.10 ✓ root matches
Cryptographically verifiable, licensed-for-AI datasets from primary sources. Every record hashed, every release signed, every diff explicit — so your retrieval layer stands on data you can defend.
free sample datasets · no card · full provenance chain
9f2c41ab…c135e41a✓ signed release"why_ragdata":
Every record carries its source URL, fetch timestamp and SHA-256. Every release ships a Merkle manifest signed via sigstore — check any record independently.
source_url · fetched_at · sha256
An explicit RAG/embedding-use license instead of a gray area, with indemnification on Pro and Enterprise. Your legal review is one page, not forty.
embedding_use: permitted
Chunk IDs are stable content addresses across releases. Pull the diff, re-embed only what changed, and keep your vector store bill flat.
+1,204 added · ~310 updated · −88 removed
US federal procurement, regulatory, courts, SEC filings, e-commerce. Collected from primary sources — never resold vendor feeds.
procurement ✓ regulatory ✓ courts ✓ sec ✓ ecommerce ✓
"how_it_works":
01 collect
Official APIs and public portals — SAM.gov, FPDS, Federal Register, EDGAR. Raw responses are snapshotted immutably and hashed at fetch time.
02 normalize
Cleaned, structured, chunked for retrieval. Chunk IDs are content addresses, so an unchanged chunk keeps its ID across releases.
03 release
JSONL plus a signed manifest per release, with diffs between versions. Verify the Merkle root before a single byte enters your index.
"catalog":
RFPs, awards and attachments from SAM.gov and FPDS — normalized NAICS/PSC codes, linked award history.
sam.gov · fpdsdaily
Federal Register and EUR-Lex with versioned diffs — track what changed, not just what exists.
federalregister.gov · eur-lexdaily
State trial court dockets normalized across fragmented portals. Personal data stripped by design.
state court portalsweekly
EDGAR filings and earnings materials, chunked for retrieval with stable section anchors.
sec.gov / edgardaily
Structured product facts for agentic commerce — attributes and availability, never copied prose.
public product feedsintraday
Missing a corpus? We build custom niches and connectors on the Enterprise plan.
contact us →
"pricing":
demo
$0
no card required
starter
$99/mo
per dataset / month
pro
$299/mo
per dataset / month
enterprise
Custom
custom agreement
"faq":
Yes — that is the point. Every dataset ships under an explicit license that permits RAG, embedding, and retrieval use. Pro and Enterprise licenses include indemnification. No scraping gray zones: we collect from primary sources, honor robots.txt and opt-outs, and carry no personal data by design.
Every record carries its source URL, fetch timestamp, and SHA-256 of the raw response. Every release ships a manifest with a Merkle root over all record hashes, signed via sigstore. Recompute any record hash, walk it up the tree, compare against the signed root — independently, with the open-source verify CLI or your own code.
Cadence is per dataset and per plan: Starter ships weekly releases, Pro ships daily. Each release is versioned with a public changelog of added, updated, and removed records, so you always know what moved.
JSONL for records and retrieval-ready chunks, plus a manifest.json per release with file hashes, counts, quality metrics, and the signed Merkle root. A Croissant (MLCommons) metadata file describes each release for dataset tooling.
Yes. Demo access is free and needs no card — sign in with your email and download sample datasets in every niche, with the full provenance chain intact. Verify first, subscribe after.
On the Enterprise plan we build custom niches and connectors against your source list, with an SLA on freshness. Tell us what your retrieval layer is missing via the contact form.