Skip to content

Repository files navigation

MasterLoop

English · 中文

You follow a math course online. You understand every lecture as you watch it. Two weeks later it's gone.

Re-watching the video is slow. Picking problems to review yourself is aimless. And re-doing a problem you've already solved is pointless — you remember the answer, not the method.

MasterLoop turns your course videos into a system that knows which problems you're about to forget, and re-skins them with fresh numbers so you actually have to solve them again — scheduled by a real spaced-repetition algorithm (FSRS), so the workload shrinks as you master things. It's a loop that converges, not an endless problem grind.

It's not flashcards you have to author. You feed in the lecture videos; it builds the knowledge graph, schedules the reviews, and regenerates the problems for you.

You drive it by talking to Codex. You don't write Python or wire scripts together — you bring your own model API keys, hand the repo to an agent (Codex / Claude Code), and tell it what you want in plain language. The agent reads the build notes + bundled skills and runs everything for you.

Alpha. The deterministic core (FSRS scheduler, SQLite store, delivery renderer) is mature and fully tested (pytest green, incl. an end-to-end smoke test). The ingest / Layer-2 / territory layers are newer — see Known limitations. Expect to iterate with the agent, not click through an installer.

Scope: math only, any math course. The concept/problem dual-track, blackboard-LaTeX extraction, and double-criterion self-check are all shaped for math. But it's any math course — exam prep, university calculus, a different teacher's lectures: feed your own. Math is expressed in LaTeX and is language-agnostic; it ships with Chinese defaults today (see the build notes for the i18n boundary).


How it works

course video
  │  ① ingest (Layer 1 + 2, needs model API keys)
  ▼
SQLite knowledge base  ── transcript / blackboard LaTeX / concept graph / problem bank / FSRS mastery / event log
  │  ② daily loop (local, no API)
  ▼
daily-practice ──► today's due problems (re-skinned) ──► you solve ──► feedback-intake (self-rate)
       ▲                                                                  │
       └──────────── FSRS reschedules, workload converges ◄───────────────┘

Three pillars:

  • Deterministic core (lib/): FSRS scheduling, SQLite knowledge base, daily delivery rendering — fully tested, subject/user-agnostic.
  • Pluggable ingest (lib/ingest/): ASR transcription + blackboard vision extraction behind a provider abstraction; swap models by editing config/providers.yaml.
  • Agent skills (.agents/skills/): daily-practice (pick + re-skin + self-check + persist), feedback-intake (parse self-rating → roll FSRS).

Who it's for

Anyone who studies from a math course and already uses a coding agent (Codex / Claude Code). You don't need to be a Python developer — the agent does the wiring. You need: a coding agent, your own model API keys (ASR / vision / LLM), and a course's lecture videos.

It is not a one-click consumer app: no install wizard, no hosted service, no GUI. The "interface" is your agent.

Quick start

The honest version of setup: get your API keys, then point your agent at this repo and talk to it.

1. Get the repo + your keys

Clone (or fork) this repo, and register for the model APIs you'll use for ingest — ASR (transcription), vision (blackboard), and an LLM (structuring + concept normalization). That key registration is the one thing only you can do; everything below, the agent handles.

2. Hand it to your agent

Open the repo in Codex (or Claude Code) and just say what you want, e.g.:

"Set this project up for me. My OpenAI/DashScope/etc. keys are in .env. The course is ; here are the lecture videos in data/lectures/. Walk me through ingesting the first lecture, then give me today's practice."

The agent reads docs/BUILDING.md + the bundled skills in .agents/skills/ and takes it from there: installs deps, fills the config files, runs ingest, builds the knowledge base, and starts the daily loop. You stay in plain language.

What the agent actually runs under the hood (you don't type these)
pip install -e ".[ingest]"                  # deps (ingest needs ffmpeg on PATH)
cp .env.template .env                        # your keys
cp config/user_preferences.yaml.template config/user_preferences.yaml
python -m lib.ingest.run --lecture-id 06 --video data/lectures/06/06.mp4 --out-dir data/lectures/06
python scripts/bootstrap_progress.py < bootstrap.json   # seed what you've already learned

Config files the agent fills for you (peek/edit anytime):

  • config/course.yaml — course identity (name / teacher / exam goal / language)
  • config/providers.yaml — which models to use for ingest
  • config/user_preferences.yaml — exam date, daily budget, studied_lectures, scheduling knobs
  • config/territory_map.yaml — concept→topic ontology (high-school-math example by default; swap for your course)

The ingest design matters: ASR and blackboard are two separate passes (the blackboard is anchored on the transcript) — the core anti-hallucination move. Layer 2 (structuring → ingestable JSON) calls your LLM through a normalize → validate → retry control loop. See the build notes.

3. Practice daily — just talk to the agent

You say The agent does
"give me today's practice" daily-practice — pick due problems, re-skin with fresh numbers, double-criterion self-check, render a PDF
"Q1 ok, Q3 forgot" / a photo of your work feedback-intake — parse your self-rating → roll FSRS, reschedule

Visualize your atlas (optional)

Your knowledge base renders as a relief map: every lecture, concept, and problem is a node, grouped into territories (the course's topics), and the terrain rises with your mastery — a cold archipelago on day one, mountain ranges where you've put in the reps. It's a WebGL2 single-canvas frontend; the layout is baked offline so it stays smooth on thousands of nodes.

MasterLoop knowledge atlas — a relief map of mastery

The terrain grows as you learn — height encodes FSRS mastery, so a topic you've drilled rises into a mountain range while untouched ground stays a low archipelago:

cold start — everything new mid mastery — ranges rising high mastery — full mountain ranges

day one (all new) · mid mastery · well-drilled — same map, terrain rising with what you've mastered

Just ask the agent: "bake and serve my knowledge atlas." It exports a snapshot, bakes the layout, and serves it locally for you.

What the agent runs
python scripts/export_snapshot.py                 # DB → frontend/data/graph.json
python scripts/bake_layout.py                     # graph.json → frontend/data/layout.json (offline force-layout)
python -m http.server 8731 --directory frontend   # then open http://127.0.0.1:8731

Re-bake after a day's review to watch the terrain grow. Positions are deterministic (same syllabus → byte-identical layout), so the map never reshuffles under you — only the mastery height changes. The baked layout.json carries your personal mastery, so it's gitignored; everyone bakes their own from their own DB.

Core contracts (don't break these)

  • No answers until you're done: problems first, solutions paginated to the end.
  • Wrong ≠ forgot: Codex never grades you; you self-rate. Arithmetic slips go to free_note, not into the FSRS signal.
  • Double-criterion self-check at generation: answer correct + main line follows the source_concepts; failing problems are discarded and replaced, never shown to you.
  • studied-set: scheduling is confined to what you've actually learned (problem-level granularity, supports partial lectures like "17:1-10").

Tests

python -m pytest        # collects tests/ only

Layout

lib/            deterministic core: fsrs/ db delivery/ migrations/ ingest/ territory.py concept_*
frontend/       WebGL2 knowledge-atlas renderer (index.html; baked layout.json is gitignored)
.agents/skills/ daily-practice / feedback-intake / concept-graph-normalize
config/         course / providers / user_preferences.template / territory_map
scripts/        bootstrap_progress / run_gate / rebuild_db / export_snapshot / bake_layout / concept_*
docs/BUILDING.md  build notes (why it's designed this way — the core deliverable)

Known limitations

Honest list (alpha; PRs welcome):

  • Layer-2 structuring needs your own LLM. lib.ingest.structurer provides the control loop (normalize + validate + retry-with-error-feedback) and a reference prompt skeleton, but no built-in LLM adapter — you pass an llm(system, user) -> str callable. Generation/structuring quality depends on your course and prompt tuning.
  • The data-quality pass needs your own LLM + agent judgment. The full dedup / merge / typo-fix / orphan-relink pipeline ships and is welded into rebuild_db --reset (it replays the ledgers deterministically). But producing the ledgers is an agent-driven adjudication step (.agents/skills/concept-graph-normalize) that calls your configured normalizer LLM — raw ingest gives you a graph with duplicates and ASR homophone typos, and the graph improves as the agent normalizes it (see docs/BUILDING.md §3).
  • territory_map.yaml is hand-authored. The default is a high-school-math example; replace it wholesale for your course. "Have an LLM generate it from your syllabus" isn't implemented yet.
  • Gemini-as-transcriber contract: Gemini is video-native and needs the video path; the default lib.ingest.run extracts audio and feeds audio.m4a, which only fits DashScope. To transcribe with Gemini, drive it directly or extend run.py (there's a warning note in the code).
  • Non-Chinese concept-ids degrade to formula-<hash> (works, but the id isn't readable); full i18n is future work.
  • Re-running a failed ingest re-does transcription (transcripts aren't cached); a failed ingest may leave an audio.m4a temp file behind.
  • ts_to_sec doesn't parse fractional seconds (12:34.56 → None).

License

MIT

About

Turn math course videos into a converging spaced-repetition loop: ingest lectures, build a concept graph, and auto-regenerate problems you're about to forget. FSRS-scheduled, LaTeX-based, language-agnostic.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages