Cerul uses model endpoints selected by the user. It does not require a Cerul account for indexing, searching, or semantic annotation. OCR runs locally on CPU with embedded weights; embedding, vision, and transcription use HTTP.
Configuration priority, from lowest to highest:
- Built-in defaults.
~/.cerul/config.toml.cerul.tomlin the current directory.CERUL_<ENDPOINT>_<FIELD>environment variables.- Repeated
--set KEY=TOML_VALUEarguments.
Multimodal embedding defaults to gemini-embedding-2, 3072 dimensions, at the
Google Gemini endpoint. A Gemini key is required for new vectors and semantic
queries. Providers, models, dimensions, and URLs remain configurable; changing
an embedding space requires compatible saved vectors or re-indexing. Offline
status, cleanup, and rebuilding cached vectors need no key.
Vision defaults to gemini-3.8-flash for explicit analysis and embodied annotation. index does not
call the vision endpoint or generate scene descriptions, chapters, or summaries.
Use cerul analyze ./video.mp4 for scenes and an overview; use
cerul annotate ./video.mp4 for embodied labels. Existing analysis and description vectors
remain available to status and search. The old --no-understanding flag is
accepted for compatibility and has no additional effect.
The CLI automatically enables the configured transcription endpoint when its key
is already available. With defaults, this reuses the Gemini key for
gemini-3.5-transcribe; indexing does not open a provider/model/credential chooser.
Use cerul config to choose Gemini, Groq, OpenAI, Custom, or Disabled. Saved keys
are reused; cerul auth set replaces the Gemini key. Explicit provider settings
and enabled = false take precedence over automatic selection.
Defaults are saved in ~/.cerul/config.toml. JSON, quiet, yes, dry-run, and redirected
invocations never prompt. Scripts and coding agents use cerul config --show to read the
effective configuration (add --json for a machine-readable object with the file path) and
cerul config --set section.field=value to save a setting without a terminal, for example
cerul config --set transcription.enabled=false or cerul config --set search.hybrid=false.
Values are TOML, so strings need quotes: --set transcription.model='"whisper-1"'.
--dry-run previews the write. Without credentials, new unconfigured ASR is skipped;
complete compatible transcripts can still be reused offline. An enabled ASR that
fails is reported as a partial result, not silently disabled.
The presets are Gemini gemini-3.5-transcribe, Groq whisper-large-v3-turbo, and
OpenAI whisper-1. Model names can be edited. Custom services need a base URL,
model, and key, and must implement OpenAI's timestamped transcription contract.
The setup check uses a silent audio clip; it checks connectivity and the response
contract, not recognition accuracy. Use a real recording to evaluate accuracy.
Keys are stored separately in ~/.cerul/credentials.json (mode 0600), scoped to
protocol, service URL, and environment variable name. Environment values override
saved keys. Configuration files contain only the environment variable name,
never the secret. cerul auth set and cerul auth remove continue to manage the
Gemini key. Library callers do not read CLI credential storage.
For non-interactive setup, export the selected key variable and configure ASR:
[embedding]
kind = "gemini"
model = "gemini-embedding-2"
dims = 3072
api_key_env = "GEMINI_API_KEY"
[vision]
kind = "openai"
base_url = "http://localhost:11434/v1"
model = "qwen3-vl"
[transcription]
enabled = true
kind = "openai"
base_url = "https://api.groq.com/openai/v1"
model = "whisper-large-v3-turbo"
api_key_env = "GROQ_API_KEY"To disable ASR permanently, set enabled = false in [transcription]. To use
Gemini, set kind = "gemini", model = "gemini-3.5-transcribe",
base_url = "https://generativelanguage.googleapis.com/v1beta", and
api_key_env = "GEMINI_API_KEY". OpenAI uses kind = "openai",
base_url = "https://api.openai.com/v1", model = "whisper-1", and
api_key_env = "OPENAI_API_KEY".
OpenAI-compatible ASR uses /audio/transcriptions with verbose_json and segment
timestamps. A model returning text without timestamps is unsupported for video
localization. Gemini's dedicated model uses native word annotations; other Gemini
models retain the prompted segment-JSON compatibility path. No Responses API or
cross-provider fallback is used. All results become the same integer-microsecond
transcript sidecar format.
Endpoint URLs cannot contain embedded credentials, query parameters, or fragments. For a command-line string override, include the TOML quotes:
cerul --set 'vision.model="qwen3-vl"' statusOrdinary status, exact text search, annotation-only filters, index rebuilds,
and cleanup do not probe providers. Commands probe only endpoints needed for
pending work. Successful probes are cached for seven days.
cerul status --providers explicitly probes embedding, vision, and enabled transcription.
Perception processing is not supported in this version. Its reserved endpoint
is probed only when explicitly configured; otherwise its capability stays null
and no request is sent to it. The command may return a partial result when one
endpoint is unavailable even if the endpoints needed for your operation work.
Use --recompute to refresh successful probes and --dry-run to avoid probes.
Capability values are true for supported, false for explicitly unsupported,
and null for unknown (including missing credentials and network errors).
A perception endpoint advertising tasks does not add processing support to this CLI.
With -v or --json, the first media request to each endpoint includes a
notice. Default human output omits routine notices; --yes suppresses the notice
in all modes. Configure endpoints before running a processing command.
Use --rpm on indexing or annotation commands to pace requests when an endpoint
returns HTTP 429. For example:
cerul index ./recording.mp4 --jobs 4 --rpm 12Choose a limit appropriate for your provider account. Cerul honors bounded
Retry-After and Gemini retry-delay hints, but retries cannot guarantee that
account quotas are available. A partial index retains completed work; repeat
the same command after quota becomes available to fill the remaining units.
Transcription uses sixty-second audio windows and converts clip-relative times
to episode-relative integer microseconds. Inspect the source when exact alignment
matters. A failed ASR does not prevent video/OCR embedding publication. Repeat
cerul index ./recording.mp4 after fixing ASR to resume missing speech windows;
completed OCR and compatible embedding checkpoints are reused. Changing the ASR
model invalidates its transcript and the derived text embeddings.
Skipping ASR omits the separate transcript. Video embedding proxies contain no
audio; speech retrieval depends on the separate transcript. JSON stream results include speech
(disabled, not_configured, no_audio, no_speech, complete, planned, or failed). A partial
failure returns exit 6 and retains a diagnostic in the stream's index state.
--workspace DIR overrides CERUL_WORKSPACE, then the default ~/.cerul.
Each embedding space includes provider kind, base URL, model, dimensions, and
query and applicable document templates in its identity. Changing any of these requires compatible new
vectors. status lists space IDs and public model metadata. A query cannot
silently search a different space.
Gemini Embedding 2's document prefix is part of the space identity. Legacy metadata without a document template keeps its original ID for inspection and rebuilding; newly prefixed documents occupy a separate space. Existing vectors are retained, and indexing into the new space requires new model embeddings.
The CLI defaults to the Gemini embedding space. Compatible endpoints and dimensions can be configured; different configurations create separate spaces. Sidecar vectors are retained by cache/index cleaning and can rebuild the corresponding index.
If you manually delete a video's .cerul directory or its episode.json, the
home page and status report missing processing data. status --json exposes
sidecar_present: false for that registration. Search and timeline views skip
it. Reading status keeps registrations intact, including paths on temporarily
disconnected drives, and never starts model requests.
To process the original video again, run cerul index /path/to/video.mp4.
Missing sidecars are recreated using the normal output location rules;
processing can call your configured models again. Deleting sidecars also removes any
annotations stored there. To forget a registration whose sidecar directory
is already gone, use cerul remove /path/to/video.mp4; the source video is kept.
Corrupt or unreadable metadata is reported as an
error with its path rather than treated as deleted data.
--json writes one final object to stdout and newline-delimited JSON events to
stderr. Human-readable mode uses stderr for progress and logs. Annotation text
is English. All annotation intervals are half-open episode-relative integer
microseconds; source PTS and clip offsets are mapped internally.
| Exit | Meaning |
|---|---|
| 0 | Success |
| 2 | Invalid arguments or configuration |
| 3 | Missing dependency or unsupported capability |
| 4 | Execution failure, including a busy workspace lock |
| 5 | Cancelled |
| 6 | Partial result; inspect incomplete stations or endpoint results |
Only one writer may operate in a workspace at a time. A lock conflict fails; commands do not silently forward to a background service. HTTP/MCP serving and perception processing are not supported in this version.
index --jobs N also bounds local OCR workers, capped by the available CPU
count. Each worker owns its model plans, so higher concurrency uses more memory.
Use --jobs 1 on memory-constrained machines. Completed frame checkpoints survive
cancellation; changing the worker count does not invalidate OCR results. Worker
completion order never changes annotation timestamps or adjacent-text merging.
OCR keeps recognized lines with confidence of at least 0.75. This suppresses low-confidence output from motion-blurred frames; it does not guarantee that every retained character is correct. The threshold is part of the station's cache identity, so results produced with an older threshold are recomputed.
When a video is replaced at the same path, Cerul registers only its current
content identity. If the adjacent sidecar belongs to the previous content, the
new sidecar uses <media>.<sha256>.cerul; the previous annotations are preserved
on disk and are not included in the current registry. Shared LeRobot video
shards continue to keep one registry entry and sidecar per episode.
[search]
hybrid = trueNormal cerul search "query" combines independent video, speech, screen, and
visual-description vectors with full-text candidates over original OCR/ASR.
No additional query embedding or generation call is made for each track.
Missing descriptions or speech simply leave those tracks absent. Image queries
use vector tracks only. Sidecars remain authoritative; both additional search
projections rebuild offline and discard stale description generations.
The initial recipe hybrid/affine-max-full-query-lexical/1 maps cosine to
0.5 * cosine + 0.5 and BM25 to 0.8 + 0.01 * BM25, clamping to [0, 1].
Only full-query token matches enter the lexical route (case-insensitive;
CJK phrases can match within unsegmented text). Video/description and
speech/screen/lexical each contribute their strongest vote, with a maximum 0.02
agreement bonus. This fixed normalization is not fitted calibration or confidence.
--threshold continues to gate raw cosine; it does not gate BM25 or the final
fused score. --text remains case-sensitive substring search.
Set hybrid = false, or pass --set search.hybrid=false, to restore three-track
raw-max ranking. Existing vectors are reused in either mode. JSON retains raw
vector scores in evidence_scores, and all admitted evidence, raw units, ranks,
original intervals, normalized values and contributing flags in fusion_evidence.
Hybrid search currently expands candidate budgets until every scoped source is
exhausted before final ranking, so agreement from lower-ranked evidence can
change the top results. Large libraries may require more time and memory; use
--in or annotation filters to narrow the scope. Each disjoint filter interval
is fused independently: later text cannot provide an excerpt or agreement bonus
to an earlier interval that it does not overlap. Evidence retains its original
source timestamps even when displayed result boundaries are clipped.