Skip to content
mnihad000Public
Public repository

Add file

Latest commit

2734b40 Â· Sep 26, 2026

History

91 Commits

Folders and files

NameName
Last commit message
Last commit date
Sep 14, 2026
Sep 26, 2026
Sep 26, 2026
Sep 26, 2026
Sep 26, 2026
Sep 22, 2026
Sep 26, 2026
Jun 21, 2026
Sep 26, 2026
Sep 26, 2026
Sep 17, 2026
Sep 17, 2026
Sep 26, 2026
Sep 22, 2026
Sep 14, 2026
Jul 16, 2026
Sep 21, 2026
Sep 26, 2026
Jul 16, 2026
Sep 21, 2026
Sep 26, 2026
Sep 15, 2026
Sep 15, 2026
Sep 15, 2026
Jun 21, 2026

Repository files navigation

RhetoriQ

RhetoriQ is an evidence-first narrative investigation system. It detects public narrative signals, retrieves source material, maps how language changes and spreads, and produces reports whose material claims point back to inspectable evidence.

The product deliberately distinguishes first observed in the available dataset from true origin and does not treat correlation as proof of coordination.

Architecture at a glance

flowchart TB
    subgraph Edge[Restricted public edge — final proof only]
        U([User]) --> B[Browser]
        DNS[Route 53 DNS] --> ALB[HTTPS ALB / Ingress]
        ACM[ACM certificate] -.- ALB
        B -->|HTTPS pages / API / SSE| ALB
        ALB -->|/| FE[React / Nginx]
        ALB -->|/api| API[FastAPI]
    end
    subgraph Core[EKS application namespace]
        API -->|accept request| PG[(PostgreSQL / pgvector)]
        PG ==>|transactional outbox| OUT[Outbox publisher]
        OUT ==>|versioned events| K[(Kafka KRaft)]
        REG[Apicurio schemas] -.- K
        K ==>|requested| IW[LangGraph investigation worker]
        K ==>|raw / enriched| FL[Flink stream job]
        FL ==>|processed / signals| DW[Document + signal workers]
        DW -->|canonical writes| PG
        IW -->|receipts + artifacts| PG
        IW --> GATE{{Claim checks + publication gate}}
        GATE -->|cited report or limitation| PG
        K ==>|projection topics| PW[Projection workers]
        PW --> ES[(Elasticsearch)]
        PW --> NG[(Neo4j)]
        PW --> V[(MiniLM vectors in pgvector)]
        API -->|validated reads| ES
        API -->|explained paths| NG
        API -->|complete-response cache| RD[(Redis cache)]
        API -->|durable history / workspace| PG
    end
    subgraph Sources[Approved acquisition]
        SX[SearXNG] --> WEB[Public records + permitted pages]
        IW -->|bounded tools| SX
        IW -->|primary APIs / fetch| WEB
    end
    classDef actor fill:#eef2ff,stroke:#4f46e5,color:#1e1b4b
    classDef service fill:#e0f2fe,stroke:#0284c7,color:#082f49
    classDef durable fill:#dcfce7,stroke:#15803d,color:#14532d
    classDef derived fill:#fef3c7,stroke:#b45309,color:#78350f
    classDef gate fill:#fce7f3,stroke:#be185d,color:#831843
    classDef trust fill:#ede9fe,stroke:#7c3aed,color:#4c1d95
    class U,B actor
    class FE,API,OUT,IW,FL,DW,PW,SX,WEB service
    class PG,K durable
    class ES,NG,V,RD derived
    class GATE gate
    class DNS,ALB,ACM,REG trust

Solid arrows show direct requests and writes; thick arrows show asynchronous events; dotted lines show trust relationships. See the full architecture for runtime processing, deployment, recovery, the complete legend, and the lifecycle walkthrough.

What the current repository implements

The repository includes the product, Kafka/Flink pipeline, B5 projections, B6 Helm chart, guarded EKS Terraform states, and deployment scripts.

  • FastAPI endpoints for ingestion, trending topics, investigations, timelines, graphs, mutations, receipts, and reports.
  • GDELT DOC 2.0 ingestion for news discovery.
  • Hacker News ingestion through the public Algolia API.
  • Direct HTTP retrieval of canonical pages for evidence enrichment.
  • A durable LangGraph research runtime with budgets, leases, checkpoints, idempotent actions, replay, SSE progress, and a deterministic publication gate.
  • Self-hosted SearXNG discovery plus GDELT, Hacker News, canonical HTTP, internal-corpus, and an isolated Playwright adapter for local research only.
  • Federal Register first-party research with policy-aware receipts, retries, pagination, deduplication, and visible limitations.
  • Apache Kafka KRaft and Apicurio Registry with 12 versioned primary topics and matching DLQs, a transactional outbox, idempotent consumers, controlled replay, and Kafka-only investigation dispatch.
  • SQLite-backed development storage and PostgreSQL/pgvector production persistence, with optional Redis capabilities.
  • A React and TypeScript investigation interface with a live graph, research rail, evidence gate, and replay controls.
  • Flink document processing and narrative signals plus B5 Elasticsearch, Neo4j, MiniLM/pgvector, and Redis projections behind acceptance-gated flags.
  • Production containers, committed PostgreSQL migrations, and CI checks for backend, frontend, schemas, documentation, and migration compatibility.
  • A shared Helm chart for kind and EKS, three isolated Terraform states, GitHub OIDC publishing to immutable ECR images, and bounded smoke, evidence, recovery, and teardown scripts.

See the documentation index, B6 operations, and roadmap for implementation and operating details.

Research strategy

RhetoriQ is agent-led and source-policy-first, not crawler-first:

  1. A user question starts a bounded investigation; the LangGraph investigator selects live broad-web search, internal-corpus recall, canonical-page retrieval, or an approved primary-source API for each evidence gap.
  2. Search results and API records are discovery leads, not automatically evidence.
  3. RhetoriQ retrieves a canonical source page when permitted and needed to create an evidence record, then preserves receipts and limitations.
  4. The event pipeline processes accepted documents and signals; additional scheduled RSS/Atom and event-stream connectors are planned for recurring monitoring.

A website, post, transcript, or official record is a source. An API, feed, or HTML fetch is the transport used to retrieve it.

Current source status is documented in DATA_SOURCES.md.

Repository layout

rhetoriq/
|-- backend/
|   |-- agents/       # planning, retrieval, synthesis, and receipts
|   |-- api/          # FastAPI route modules
|   |-- models/       # shared Pydantic contracts
|   |-- services/     # ingestion, retrieval, analysis, and persistence
|   `-- tests/
|-- frontend/         # React, TypeScript, and Vite application
|-- docs/             # design and operating documentation
|-- deploy/helm/      # kind and EKS application chart
|-- infra/terraform/  # isolated AWS bootstrap, foundation, and platform states
|-- SYSTEM_DESIGN.md
`-- README.md

Local development

Backend

cd backend
python -m venv ..\.venv
..\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
uvicorn main:app --reload

The API is available at http://127.0.0.1:8000; interactive documentation is at /docs.

The running OpenAPI document is the authoritative endpoint reference. The API includes health, ingestion, narratives, trending, investigations, research events/replay, evidence search, and provenance-path routes. Accepted ingestion and investigation work is Kafka-only; the API commits an outbox record and returns without executing the job synchronously.

Frontend

cd frontend
npm install
npm run dev

Vite normally serves http://127.0.0.1:5173. Routes are /, /dashboard, and /investigation/:id. The production Nginx image reads PUBLIC_API_BASE_URL at container startup; local Vite development uses VITE_API_BASE_URL.

Full local stack

The root Compose stack supplies PostgreSQL/pgvector, Kafka, Apicurio Registry, SearXNG, topic initialization, the outbox publisher, role-scoped workers, Flink, API, and frontend:

$env:POSTGRES_PASSWORD="<local-secret>"
$env:SEARXNG_SECRET="<local-secret>"
docker compose up --build -d
docker compose ps

See Architecture for research and service boundaries and Operations for startup, replay, and recovery.

Tests

pytest backend/tests
cd frontend
npm run build

Important configuration

Settings are loaded from backend/.env when present.

Variable Purpose
DEMO_MODE Use the bundled demo corpus and skip background live refreshes.
GEMINI_API_KEY Optional Gemini model access.
GROQ_API_KEY Optional Groq model access.
RESEARCH_RUNTIME auto, native, or langgraph runtime selection.
RESEARCH_EXECUTION_MODE Deployment label; execution is Kafka-only and defaults to kafka.
KAFKA_BOOTSTRAP_SERVERS Kafka bootstrap addresses.
KAFKA_SCHEMA_REGISTRY_URL Apicurio Confluent-compatible API base URL.
KAFKA_CLIENT_ID / KAFKA_CONSUMER_GROUP_PREFIX Producer and consumer identity prefixes.
KAFKA_SECURITY_PROTOCOL Kafka transport security protocol.
KAFKA_SASL_USERNAME / KAFKA_SASL_PASSWORD Optional SASL credentials; never written to events or logs.
KAFKA_TOPIC_PREFIX Optional environment-specific physical topic prefix.
KAFKA_RETRY_MAX_ATTEMPTS / KAFKA_RETRY_BACKOFF_SECONDS Bounded consumer/producer retry policy.
SEARXNG_BASE_URL Self-hosted broad-search endpoint.
BROWSER_SERVICE_URL Isolated browser-rendering endpoint.
BROWSER_RENDERING_ENABLED Keep false in the initial public deployment; enables the local browser adapter only when explicitly configured.
DATABASE_URL PostgreSQL connection string for production persistence; use the managed Neon value and keep sslmode=require when supplied by Neon.
DEPLOYMENT_ENV Set to production on the Railway API service; production startup requires DATABASE_URL.
CORS_ALLOW_ORIGINS Comma-separated public frontend origin(s) allowed by the API.
ENABLE_POSTGRES_VECTOR_SEARCH Feature flag for the additive Neon pgvector retrieval path; keep false until migration, backfill, and comparison checks pass.
POSTGRES_VECTOR_SEARCH_TOP_K Maximum persisted semantic corpus results when the Neon path is enabled.
POSTGRES_VECTOR_BACKFILL_BATCH_SIZE Resumable Neon corpus backfill batch size.
REDIS_URL Optional Redis cache, phrase store, vector store, and memory.
GDELT_BASE_URL GDELT DOC 2.0 endpoint.
GDELT_MAX_RECORDS Maximum GDELT records requested per query.
FETCH_TIMEOUT_SECONDS Canonical-page retrieval timeout.

Connector credentials are configured only for approved deployments. Reddit access, commercial search APIs, and licensed news products require terms and retention review before production use.

Documentation

Document Purpose
SYSTEM_DESIGN.md Concise system principles and investigation lifecycle.
Documentation index Entry point for all durable project documentation.
ARCHITECTURE.md End-to-end architecture, runtime events, deployment, trust, recovery, and lifecycle.
DATA_SOURCES.md Source hierarchy, provider status, and compliance requirements.
KAFKA.md Implemented replayable event contracts and operations.
OPERATIONS.md Startup, health, replay, recovery, projection administration, and troubleshooting.
DEPLOYMENT.md Public release, rollback, Kubernetes boundary, and AWS demonstration strategy.
RELEASE_READINESS.md Current candidate evidence and the ordered gates remaining before AWS.
PRE_B6_GUIDE.md Current workstation progress and readiness sequence before Kubernetes.
B6_OPERATIONS.md Implemented full-stack Helm/EKS runbook, current blockers, evidence gates, and teardown.
TESTING.md Routine checks and B3–B5 acceptance gates.
100K stress-test plan Standalone daily-capacity experiment.
ROADMAP.md Delivery status, remaining acceptance, and future phases.