- Keep
mainand thecheckpoint/0X-*tags consistent for implementation changes. When changing workshop code or docs, inspect and update any checkpoint tags that contain the same affected state before considering the task done.
- Keep the root
README.mdfocused on learners and instructors: workshop scope, entry points, module map, and minimal checkpoint usage. - Do not recreate overview files such as
docs/README.md,docs/learner/README.md, ordocs/checkpoints.md. Put learner-facing overview content in the root README or the learner lessons; put facilitator content in instructor lessons; put maintainer-only checkpoint strategy here.
The workshop needs two things at once:
- A clean linear story that matches the Langfuse AI engineering loop.
- The ability to jump ahead when a live workshop runs out of time.
Maintain this repo strategy:
- Keep
mainas the complete reference app plus current docs. - Keep
checkpoint/00-setupequivalent tocheckpoint/01-base-appfor environment validation. - Keep one milestone tag for each workshop step.
- Make every later step runnable through explicit fallbacks.
Canonical checkpoint tags:
checkpoint/00-setupcheckpoint/01-base-appcheckpoint/02-tracingcheckpoint/03-prompt-managementcheckpoint/04-monitoringcheckpoint/05-datasetcheckpoint/06-experimentscheckpoint/07-evaluationcheckpoint/08-wrap-up
Canonical progression:
00-setup: setup, keys, Langfuse Cloud EU, Langfuse CLI, Langfuse skill, and workshop framing on the same untraced app state as01-base-app.01-base-app: working Dad IT Support Agent on the official OpenAI SDK, with one fixed Dad context, two local tools, and no Langfuse tracing yet.02-tracing: learners add Langfuse tracing on top of the base app with OpenTelemetry setup,observeOpenAI(new OpenAI()),observe(...)wrappers around app/tool functions, and optionalpropagateAttributes(...)for user/session metadata.03-prompt-management: learners replace the code-only prompt path with a Langfuse-managed prompt plus local fallback.04-monitoring: starts from the traced app and stable message-array trace shape; this step is mostly evaluator design, variable mapping, and Langfuse UI setup.05-dataset: adds a starter dataset that matches the app scope and uses message-array inputs plus expected outputs.06-experiments: runs the app against the Langfuse dataset with the SDK experiment runner and tworunExperimentcallback evaluators — deterministickeyword_overlapplus an in-script LLM-as-a-judgecorrectnessscore. Platform evaluators are optional bonus only (they need existing experiment data to configure).07-evaluation: changes the prompt, reruns the same dataset, and compares the samekeyword_overlap+correctnesscallback scores side by side.08-wrap-up: recaps the mental model and points to next steps.
What makes the checkpoints stitchable:
- The OpenAI SDK is already in place before tracing starts, so step 2 is only about observability.
- Prompt management falls back to the local prompt if the Langfuse prompt is absent.
- Monitoring depends on stable message arrays on the agent/generation observations and the root
answerfield, not provider-specific internals. - Dataset items carry both
idealAnswerandexpectedKeywords, which feed the step-06correctnessandkeyword_overlapcallback evaluators. - Dataset and experiment scripts reuse the same app logic as the web UI; experiment scores are produced by
runExperimentcallbacks in the runner process.
Recommended jump patterns:
- Short workshop: start at
checkpoint/01-base-app, build tracing live, explain prompt management, and finish with monitoring. - Full workshop: walk through all checkpoints in order.
- Catch-up jump: if a group gets stuck in tracing, jump straight to
checkpoint/04-monitoringorcheckpoint/05-datasetand continue from there.