Skip to content
sh4mbhaviPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

 ▄████▄   ▄▄▄        ██████  ▄████▄   ▄▄▄      ▓█████▄ ▓█████ 
▒██▀ ▀█  ▒████▄    ▒██    ▒ ▒██▀ ▀█  ▒████▄    ▒██▀ ██▌▓█   ▀ 
▒▓█    ▄ ▒██  ▀█▄  ░ ▓██▄   ▒▓█    ▄ ▒██  ▀█▄  ░██   █▌▒███   
▒▓▓▄ ▄██▒░██▄▄▄▄██   ▒   ██▒▒▓▓▄ ▄██▒░██▄▄▄▄██ ░▓█▄   ▌▒▓█  ▄ 
▒ ▓███▀ ░ ▓█   ▓██▒▒██████▒▒▒ ▓███▀ ░ ▓█   ▓██▒░▒████▓ ░▒████▒
░ ░▒ ▒  ░ ▒▒   ▓▒█░▒ ▒▓▒ ▒ ░░ ░▒ ▒  ░ ▒▒   ▓▒█░ ▒▒▓  ▒ ░░ ▒░ ░
  ░  ▒     ▒   ▒▒ ░░ ░▒  ░ ░  ░  ▒     ▒   ▒▒ ░ ░ ▒  ▒  ░ ░  ░
░          ░   ▒   ░  ░  ░  ░          ░   ▒    ░ ░  ░    ░   
░ ░            ░  ░      ░  ░ ░            ░  ░   ░       ░  ░
░                           ░                   ░             

Contextual Analysis of Spread for Content Authenticity Detection and Evaluation

Tests Python mypy strict Licence: MIT

A forensic CLI tool that scores Reddit propagation patterns to flag suspected deepfakes and produces a chain-of-custody PDF report.


What it does

CASCADE ignores the pixels and looks at how content spreads on Reddit. Authentic media tends to land in a topically coherent subreddit, get engaged with by regulars, and crosspost in a sensible order. Synthetic or deepfake-style content dropped in cold often misses this pattern: unusual subreddit, new or suspicious posting account, off-key engagement, no crossposts or strange ones.

CASCADE turns that observation into six measurable features, scores them with a calibrated Random Forest, and outputs a forensic-grade PDF with full chain of custody — append-only audit log, hash manifest, perceptual hash for repost detection, and a single master hash that binds the substantive evidence into a tamper-detectable signature.

The six features

F# Feature What it measures
F1 account_age_days Days between posting account creation and post date
F2 account_karma Total Reddit karma at time of posting
F3 subreddit_coherence TF-IDF cosine similarity between post text and subreddit's reference vocabulary
F4 time_to_first_crosspost_s Seconds to first observed crosspost (censored at 7-day window)
F5 first_hour_comment_velocity Distinct non-deleted comments in the first 3,600 seconds
F6 network_distance_mean Mean shortest-path distance among first three engagers in a co-engagement graph

Headline result

H1: STRONG SUPPORT in both dataset variants (60-item simulated dataset; LOOCV with 10,000-resample bootstrap CI; isotonic calibration).

Variant AUC-ROC (LOOCV) 95% CI
hybrid 0.9922 [0.9744, 1.0000]
independent 0.9856 [0.9577, 1.0000]

Results on simulated data demonstrate that the CASCADE methodology is sound in principle. They do not demonstrate that real-world deepfakes exhibit these propagation patterns at the rate the simulation assumes. Real-world validation is explicit future work.


What it does not do

  • It does not look at content itself. No pixel analysis, no face detection, no audio fingerprinting.
  • It does not produce a verdict. Output is P(synthetic) with a 95% confidence interval.
  • It is not court-admissible by default. It produces evidence that can support an admissibility argument — admissibility is the court's call.
  • It was not validated on real-world deepfake cases. Real-world validation is explicit future work.

Install

Option A — pip install from GitHub (recommended)

No dataset or model training needed. The trained classifier is bundled.

pip install git+https://github.com/sh4mbhavi/cascade.git
cascade

Option B — clone and run locally

Requires Python 3.11.

git clone https://github.com/sh4mbhavi/cascade.git
cd cascade
python3.11 -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -r requirements.txt
make test

Usage

Automatic collection (recommended)

Provide a Reddit URL and your image. CASCADE fetches everything else — no account or API key needed.

cascade collect \
  --url "https://www.reddit.com/r/worldnews/comments/abc123/post_title/" \
  --media ./image.png \
  --case MY-CASE-001

cascade analyse --case MY-CASE-001 --seed 42
cascade verify  --case MY-CASE-001

Manual collection

If the post is deleted or the subreddit is private, build propagation.json yourself. See docs/COLLECT.md for the full guide and schema.

Output

Each run produces cases/<case-id>/:

File What it is
report.pdf Forensic PDF — score, feature breakdown, propagation graph, chain of custody
audit.jsonl Append-only audit log — every stage, every hash, every timestamp
manifest.json SHA-256 manifest of all artefacts, bound by the master hash
artefacts/ Graph render, feature importance chart, GraphML

Quick example

make example          # generates cases/EXAMPLE/ with a sample case
make verify-example   # verifies chain of custody — all green

Tamper detection:

echo "tampered" >> cases/EXAMPLE/input/media.png
make verify-example   # exits non-zero, shows exactly what changed
make example          # regenerates clean

Architecture

Five-stage pipeline: ingest → graph → features → score → report

Each stage writes append-only entries to audit.jsonl and contributes hashed artefacts to manifest.json. The score stage runs a calibrated Random Forest and emits P(synthetic) with a 95% bootstrap CI. The report stage renders a ReportLab PDF that embeds the master hash on the cover page, every-page footer, and appendix.

The master hash uses Z'-extended semantics — it excludes report.pdf, manifest.json, and audit.jsonl from the canonical hash so the binding can be embedded inside the PDF without a circular dependency, while all three files remain in manifest.files for per-file tamper detection.


Reproduce the classifier

make simulate    # regenerates the 60-item dataset
make train       # trains both variants, writes to cascade/models/

Testing

make test

153 tests across 11 modules. CI runs the same suite plus ruff, black --check, and mypy --strict on every push.


Licence

MIT. CASCADE outputs are not court-admissible by themselves; users intending forensic use should obtain independent expert review and validate against their jurisdiction's evidentiary requirements.


built by Sham

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages