Back to Projects
RhetoriQ cover

2026 · ACTIVE DEVELOPMENT

Trace How Narratives Spread

An evidence-first investigation system that follows public narratives through source material, competing claims, and inspectable receipts.

Overview

I built a FastAPI and React/TypeScript product for investigating how public narratives emerge and change. The workspace brings together live research progress, a report, source evidence, a timeline, a narrative graph, and a claim-level audit. It distinguishes the first observation in available data from a proven origin and does not infer coordination from correlation alone.

The LangGraph research runtime plans bounded investigations, selects approved sources, retrieves canonical pages, saves provenance receipts, and resumes from durable checkpoints. A verification layer checks exact evidence spans, source independence, entailment, and contradictions before a publication gate allows supported claims into a report; weak or missing evidence is shown as a limitation.

The event backbone uses Kafka KRaft, Apicurio schemas, a transactional outbox, role-scoped consumers, retries, dead-letter queues, and replay. A Flink job implements stateful document processing and narrative signals. PostgreSQL is the authoritative store; Elasticsearch, Neo4j, MiniLM/pgvector, and Redis provide recoverable search, graph, semantic, and cache projections.

I also built Docker/Compose packaging, a shared Helm chart for kind and EKS, three guarded Terraform states, immutable-image delivery, and scripts for smoke tests, recovery, evidence capture, and teardown. The repository includes static validation and local regression results; public launch and full actual-stack qualification remain underway.

Implemented product and infrastructure: 386 backend tests and 27 frontend tests passed in the recorded local release regression. Public URLs, EKS deployment, and the remaining B3–B5 acceptance gates are still pending.

From the Repository

README / Project overview

Trace every claim to its source.

RhetoriQ detects public narrative signals, retrieves source material, maps how language changes and spreads, and produces reports whose material claims point back to inspectable evidence.

The system distinguishes the first observation in available data from a proven origin, and it does not treat correlation as proof of coordination.

01

Research with receipts

Bounded LangGraph investigations use approved public sources, canonical retrieval, and provenance records.

02

Claims checked against evidence

Exact spans, source independence, and contradictions inform what a report can publish.

03

A replayable pipeline

Kafka events, Flink processing, PostgreSQL authority, and recoverable search and graph projections carry the data.

How It Works

1) A user submits a question or ingests a source document.
2) FastAPI saves the request and a transactional outbox record in PostgreSQL.
3) The outbox publisher sends a versioned event to Kafka; workers consume with idempotency and replay controls.
4) LangGraph plans bounded research and chooses SearXNG, Federal Register, GDELT, Hacker News, internal recall, or canonical-page retrieval.
5) Every usable source receives a provenance receipt. Flink processes documents and computes narrative signals.
6) Claim checks inspect evidence spans, independence, contradictions, and missing support.
7) The publication gate writes a cited report or a visible limitation to PostgreSQL.
8) React displays live SSE progress, the report, evidence library, timeline, and graph; search and graph projections enrich the workspace.

Architecture

rhetoriq/
  frontend/             React + TypeScript investigation workspace
  backend/
    api/                 FastAPI ingestion, investigations, search, graph, SSE
    agents/              LangGraph planning, retrieval, receipts, verification
    services/            persistence, events, analysis, projections
    tests/               backend and contract coverage
  infra/
    flink/               stateful document and narrative-signal processing
    research/            SearXNG and isolated browser adapter
    terraform/eks-demo/  bootstrap, foundation, platform states
  deploy/helm/rhetoriq/  shared kind and EKS chart
  docs/                   architecture, operations, testing, roadmap

  PostgreSQL -> outbox -> Kafka -> workers/Flink -> PostgreSQL
                                         -> ES / Neo4j / pgvector / Redis

Datasets

GDELT DOC 2.0

News discovery and narrative leads, followed by canonical-source retrieval when evidence is needed.

Open dataset

Hacker News Algolia API

Public discussion discovery through the implemented Algolia ingestion adapter.

Open dataset

Federal Register

First-party policy records with pagination, retries, provenance receipts, and explicit limitations.

Open dataset

Approved public web and internal corpus

SearXNG finds leads; policy-aware canonical retrieval and persisted documents supply inspectable evidence.

Open dataset

Setup

Prerequisites

  • Docker Desktop
  • Python
  • Node.js
  • Credentials for optional live providers

Installation

git clone https://github.com/mnihad000/rhetoriq.git
cd rhetoriq
# Configure the local secrets described in README.md.
docker compose up --build -d

Environment

POSTGRES_PASSWORD and SEARXNG_SECRET are required by the local Compose stack.
Optional: GEMINI_API_KEY or GROQ_API_KEY for hosted model access.
Production uses DATABASE_URL and deployment-specific Kafka, CORS, and research settings.
See README.md and docs/OPERATIONS.md for the full configuration.

Connect Services

Frontend: http://127.0.0.1:5173 in Vite development
API: FastAPI routes under /api
Investigation workspace: /investigation/:id

Model Setup

The MiniLM semantic projection uses sentence-transformers/all-MiniLM-L6-v2.
Claim verification uses a local NLI model when configured.
See docs/OPERATIONS.md for model and feature-flag setup.

Run Services

docker compose ps
pytest backend/tests
cd frontend
npm run build

Decision Engine

A user question starts a bounded, checkpointed investigation. The worker selects source adapters for each evidence gap, records receipts, and passes proposed claims through deterministic publication rules. Kafka and Flink process ingestion and narrative signals asynchronously; retries and replay preserve progress when workers restart.

State Snapshot (Input)

POST /api/investigate
{"query_text":"How did this public claim spread?"}

Structured Action (Output)

{
  "investigation_id": "inv_<generated-id>",
  "status": "planning_completed",
  "current_stage": "planner",
  "query_text": "How did this public claim spread?",
  "plan": { "...": "bounded research plan" },
  "warnings": []
}

Decision Triggers

  • A submitted question or ingested document creates durable investigation or processing work.
  • Research gaps determine which approved source adapter the LangGraph worker calls next.
  • Verified evidence, contradictions, and independence checks determine whether a claim is published, qualified, or withheld.

Source & Provenance Analysis

  • Receipts keep canonical URLs, retrieval context, source roles, and evidence spans inspectable.
  • Timeline and graph views show observed paths and changes in language without treating an observed first source as the true origin.
  • Elasticsearch, Neo4j, and pgvector projections are revalidated against PostgreSQL before results are shown.