▄████▄ ▄▄▄ ██████ ▄████▄ ▄▄▄ ▓█████▄ ▓█████
▒██▀ ▀█ ▒████▄ ▒██ ▒ ▒██▀ ▀█ ▒████▄ ▒██▀ ██▌▓█ ▀
▒▓█ ▄ ▒██ ▀█▄ ░ ▓██▄ ▒▓█ ▄ ▒██ ▀█▄ ░██ █▌▒███
▒▓▓▄ ▄██▒░██▄▄▄▄██ ▒ ██▒▒▓▓▄ ▄██▒░██▄▄▄▄██ ░▓█▄ ▌▒▓█ ▄
▒ ▓███▀ ░ ▓█ ▓██▒▒██████▒▒▒ ▓███▀ ░ ▓█ ▓██▒░▒████▓ ░▒████▒
░ ░▒ ▒ ░ ▒▒ ▓▒█░▒ ▒▓▒ ▒ ░░ ░▒ ▒ ░ ▒▒ ▓▒█░ ▒▒▓ ▒ ░░ ▒░ ░
░ ▒ ▒ ▒▒ ░░ ░▒ ░ ░ ░ ▒ ▒ ▒▒ ░ ░ ▒ ▒ ░ ░ ░
░ ░ ▒ ░ ░ ░ ░ ░ ▒ ░ ░ ░ ░
░ ░ ░ ░ ░ ░ ░ ░ ░ ░ ░ ░
░ ░ ░
Contextual Analysis of Spread for Content Authenticity Detection and Evaluation
A forensic CLI tool that scores Reddit propagation patterns to flag suspected deepfakes and produces a chain-of-custody PDF report.
CASCADE ignores the pixels and looks at how content spreads on Reddit. Authentic media tends to land in a topically coherent subreddit, get engaged with by regulars, and crosspost in a sensible order. Synthetic or deepfake-style content dropped in cold often misses this pattern: unusual subreddit, new or suspicious posting account, off-key engagement, no crossposts or strange ones.
CASCADE turns that observation into six measurable features, scores them with a calibrated Random Forest, and outputs a forensic-grade PDF with full chain of custody — append-only audit log, hash manifest, perceptual hash for repost detection, and a single master hash that binds the substantive evidence into a tamper-detectable signature.
| F# | Feature | What it measures |
|---|---|---|
| F1 | account_age_days |
Days between posting account creation and post date |
| F2 | account_karma |
Total Reddit karma at time of posting |
| F3 | subreddit_coherence |
TF-IDF cosine similarity between post text and subreddit's reference vocabulary |
| F4 | time_to_first_crosspost_s |
Seconds to first observed crosspost (censored at 7-day window) |
| F5 | first_hour_comment_velocity |
Distinct non-deleted comments in the first 3,600 seconds |
| F6 | network_distance_mean |
Mean shortest-path distance among first three engagers in a co-engagement graph |
H1: STRONG SUPPORT in both dataset variants (60-item simulated dataset; LOOCV with 10,000-resample bootstrap CI; isotonic calibration).
| Variant | AUC-ROC (LOOCV) | 95% CI |
|---|---|---|
| hybrid | 0.9922 | [0.9744, 1.0000] |
| independent | 0.9856 | [0.9577, 1.0000] |
Results on simulated data demonstrate that the CASCADE methodology is sound in principle. They do not demonstrate that real-world deepfakes exhibit these propagation patterns at the rate the simulation assumes. Real-world validation is explicit future work.
- It does not look at content itself. No pixel analysis, no face detection, no audio fingerprinting.
- It does not produce a verdict. Output is
P(synthetic)with a 95% confidence interval. - It is not court-admissible by default. It produces evidence that can support an admissibility argument — admissibility is the court's call.
- It was not validated on real-world deepfake cases. Real-world validation is explicit future work.
No dataset or model training needed. The trained classifier is bundled.
pip install git+https://github.com/sh4mbhavi/cascade.git
cascadeRequires Python 3.11.
git clone https://github.com/sh4mbhavi/cascade.git
cd cascade
python3.11 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
make testProvide a Reddit URL and your image. CASCADE fetches everything else — no account or API key needed.
cascade collect \
--url "https://www.reddit.com/r/worldnews/comments/abc123/post_title/" \
--media ./image.png \
--case MY-CASE-001
cascade analyse --case MY-CASE-001 --seed 42
cascade verify --case MY-CASE-001If the post is deleted or the subreddit is private, build propagation.json yourself. See docs/COLLECT.md for the full guide and schema.
Each run produces cases/<case-id>/:
| File | What it is |
|---|---|
report.pdf |
Forensic PDF — score, feature breakdown, propagation graph, chain of custody |
audit.jsonl |
Append-only audit log — every stage, every hash, every timestamp |
manifest.json |
SHA-256 manifest of all artefacts, bound by the master hash |
artefacts/ |
Graph render, feature importance chart, GraphML |
make example # generates cases/EXAMPLE/ with a sample case
make verify-example # verifies chain of custody — all greenTamper detection:
echo "tampered" >> cases/EXAMPLE/input/media.png
make verify-example # exits non-zero, shows exactly what changed
make example # regenerates cleanFive-stage pipeline: ingest → graph → features → score → report
Each stage writes append-only entries to audit.jsonl and contributes hashed artefacts to manifest.json. The score stage runs a calibrated Random Forest and emits P(synthetic) with a 95% bootstrap CI. The report stage renders a ReportLab PDF that embeds the master hash on the cover page, every-page footer, and appendix.
The master hash uses Z'-extended semantics — it excludes report.pdf, manifest.json, and audit.jsonl from the canonical hash so the binding can be embedded inside the PDF without a circular dependency, while all three files remain in manifest.files for per-file tamper detection.
make simulate # regenerates the 60-item dataset
make train # trains both variants, writes to cascade/models/make test153 tests across 11 modules. CI runs the same suite plus ruff, black --check, and mypy --strict on every push.
MIT. CASCADE outputs are not court-admissible by themselves; users intending forensic use should obtain independent expert review and validate against their jurisdiction's evidentiary requirements.
built by Sham