Skip to content

ARTHR: the purpose-built engine behind every alert we ship.

Advanced Regulatory Tracking, Harvesting and Response. The pipeline that turns regulator websites into typed, queryable records — and the engine that makes SWORD’s data better than anything you could build, buy, or scrape elsewhere. Built from scratch by regulatory experts, not retrofitted from a generic AI platform or a content farm.

Book a call

From regulator websites to your stack

CONTINUOUS MONITORING. Scrapers running against every regulator in coverage, around the clock. New publications detected as they're posted, not on a nightly batch.

AI EXTRACTION AND CLASSIFICATION. Every alert extracted, normalized, and classified against a shared taxonomy of sectors, jurisdictions, and alert types. The classification logic is built by regulatory experts not generic NLP guessing from keywords.

LINKED, STRUCTURED RECORDS. Every alert tied to the regulator that issued it, the rule it implements, the jurisdictions it affects, and the prior alerts it supersedes. Not a flat keyword index — a graph your product can traverse.

RegAlytics Branding

Why we built ARTHR from scratch.

ARTHR is the rebuild for the world of AI, not just an iteration.

RegAlytics® Classic ran for six years. We grew it into the largest regulatory dataset on the market.

Then we tore it down.

Not because it failed – because customers started doing things with regulatory data we couldn’t have predicted in 2019. Feeding LLMs. Training models. Building agents. Embedding regulatory data into products their own customers never see. The old foundation was built for a world that no longer exists.

We chose to rebuild the entire operational stack – not just what customers touch, but everything behind it. Twelve months. Three production applications. A regulatory dataset purpose-built for AI workloads. The pieces behind the curtain are as engineered as the ones in front of it.

ARTHR is the engine. SWORD is what your team builds with.

Source and interpretation, kept structurally separate

Most regulatory data products mix what the regulator published with what someone else concluded about it. The headline summary, the keyword tags, the issuing agency, the publication date — all collapsed into one flat record, with no way to tell which fields came from the source and which came from the vendor’s analysis.

SWORD doesn’t work that way. Every alert is built from two top-level sections: what the regulator said (the source document, captured faithfully) and what RegAlytics concluded (the structured analysis layered on top). Each lives in its own part of the record, with its own fields, its own sourcing, and its own auditability. The distinction is structural: the document side holds extractions – things pulled directly from the source – and the analysis side holds enrichments – things RegAlytics generated. Different epistemological category, different rules.

That separation isn’t aesthetic. It’s the difference between a dataset you can defend in a regulatory audit and one you can’t. Between an AI agent that cites a primary source and one that confidently hallucinates a paraphrase. Between trusting an interpretation because you can verify it against the original and trusting it because the vendor said so.

The structure is the integrity. You can always read what was published. You can always inspect what was concluded. You never have to take one as the other.

The shape of an alert

Every alert SWORD ships has three top-level sections. Each does a different job. Each is queryable independently.

document

What the regulator published

Captured faithfully from the source: the body text, the publishing entity, the publication date, the document type, and the URL to the original. No editorial interpretation, no inferred fields. If a downstream system needs to verify or cite the original, this is the section it reads.

Fields include: publishers, title, url, dates, alert_type, source_language, citations, events, identified.

analysis

What RegAlytics® concluded

The structured intelligence layered on top: every classification is scored 0–1.0 with a written justification, not a binary yes/no. Sectors, alert classification, relevance – each carries a score and an audit trail explaining how ARTHR reached it. Filter at any threshold that fits your use case: 80% by default, lower if your product needs broader recall, higher if it needs precision.

Fields include: classification.categories (alert type scoring), classification.sectors (industry sector scoring), summary.

meta

Alert lifecycle

The record-keeping layer: when the alert was ingested, when it was last updated, the version number, which classifications were run, and the audit trail of any changes. If a downstream system needs to determine recency, reconcile state, or replay history, this is the section it reads.

Fields include: record.version, record.sourced_at, record.last_modified_at, record.change_type, pipeline.processing_status, pipeline.stages, pipeline.source.

Designed for how you build

Most regulatory data products are stored well and queried poorly. SWORD is built for the way modern products actually consume data — three query patterns, one consistent contract underneath. Use one, combine all three, layer them into a single request.

Content matching

Match against the document body — either literally (every alert that says “third-party risk management“) or semantically (alerts about AI training data, even when the language doesn’t overlap). Two query modes, one indexed corpus. For users searching by keyword. For AI agents reasoning about meaning. For one query that needs both.

Structured filtering

Query against the classification layer directly. Filter by sector, jurisdiction, publisher, alert type, date range – combined with boolean logic. For when your product needs precision, not relevance.

Threshold-based filtering

Classifications are scored, not boolean. Default at 80% confidence – adjust per-query, per-field, per-use-case. Filter for “alerts at least 40% financial” if your product needs broad recall, or “95%+ enforcement matters” if it needs surgical precision. The threshold is your call, not ours.

RegAlytics Branding

Connecting related alerts

Regulators don’t publish in clean atomic units. They publish proceedings — multi-step regulatory events that unfold over months or years — and omnibus releases that bundle dozens of unrelated rule changes into a single document. Most data products either keep these together as one blob (impossible to filter), or split them apart and lose the connection (impossible to understand). SWORD does both jobs at once.

Proceedings
A proposed rule, the public comment period, the final rule, the enforcement actions taken under it — those aren’t four unrelated alerts. They’re one regulatory story unfolding over time, and the connections between them carry real information.

Where regulators publish with shared citation references, SWORD links those alerts together. Pull the citation, see every alert tied to it. Subscribe to a citation, receive every subsequent update against it. The linking is conservative — we only connect what the data supports — but where it exists, it works exactly the way you’d want.

Omnibus decomposition
A single congressional bill might modify forty separate statutes across twelve agencies. A single regulator publication might cover banking, insurance, and consumer protection in one document. Treated as one alert, it’s noise. Split blindly into many, it loses the relationships.

SWORD decomposes omnibus releases into their distinct components — each tagged with the right sector, jurisdiction, and classification — while preserving the link back to the original publication. Your code queries by topic; your audit trail still resolves to the source.

Built for production

The data layer that sits underneath a real product needs to behave like infrastructure — predictable, recoverable, and forward-compatible. SWORD is designed against those expectations from the ground up.

Versioned and replayable

Every alert carries a version number that increments on update. Re-pull any window of history, reconcile against your own store, or replay an event stream from any point. State recovery doesn’t require a support ticket.

Data quality is shared

When something looks wrong — a miscategorized alert, a stale field, a duplicate — the fix flows back into the same record. Versioning means corrections don’t break your downstream state; they just append. You see what changed and when, and your replays stay deterministic.

Designed to grow

The schema absorbs new fields, classifications, and extraction types without restructuring — when a new alert variant emerges (say, a fine type unique to one jurisdiction), it lands in the right sub-block with the right label. If your code pulls that sub-block, you pick up the new content automatically. Additive changes only; your existing integration doesn’t break when we ship new capability.

RegAlytics Branding

Put SWORD under your stack.

Thirty minutes. Bring your architecture, your use case, and the question the next page of your roadmap depends on. We’ll talk through how SWORD fits, what the integration looks like, and where the actual gains show up.

Name(Required)