→ posturi.gov2.ro [WIP]
Alternative browser / explorer for posturi.gov.ro. Scrapes the Romanian government job listings portal — and tracks changes over time. Pipeline: index → cache announcement pages → extract structured data → LLM-extracted structured display sections stored in Postgres.
Derivative work: mariuscomper.uk/posturi-publice
The October 3 project audit records the code, documentation and live-site findings. Use the consolidated backlog for current priorities and the remediation specifications for bounded coding-agent assignments. Historical tracking is preserved in the activity log and the pre-consolidation backlog.
The agreed product direction captures the user-facing feature scope and recommended sequence toward an application companion.
flowchart LR
web(["posturi.gov.ro"])
subgraph scrape["① Scrape"]
direction TB
fetchIdx["fetch-index.py"]
indexCSV[/"posturi_gov_ro.csv"/]
fetchDetail["fetch-anunturi.py"]
htmlCache[/"anunturi/\n**∕*.html"/]
download["download-attachments.py"]
dlFiles[/"downloads/\n*.docx *.pdf"/]
fetchIdx --> indexCSV --> fetchDetail --> htmlCache
htmlCache --> download --> dlFiles
end
subgraph parse["② Parse"]
direction TB
parseScript["parse-anunturi.py"]
anunturiCSV[/"anunturi.csv"/]
calendarCSV[/"calendar.csv"/]
parseScript --> anunturiCSV & calendarCSV
end
subgraph db_layer["③ Import → Postgres"]
direction TB
importCSVs["import_csvs"]
extractCmd["extract_attachments"]
inferCmd["infer_postings"]
pg[("jobs_jobposting\nbody_markdown\nattachment_text\ninferred JSONB")]
importCSVs --> pg
extractCmd -->|"attachment_text"| pg
pg --> inferCmd -->|"inferred"| pg
end
subgraph enrich["④ Enrich"]
llmSchema["llm-schema.py"]
end
subgraph serve["⑤ Serve"]
direction TB
pg[/"PostgreSQL"/]
sqliteExport["export-to-sqlite.py\n--active-only"]
sqliteDB[/"posturi.sqlite\n(active only)"/]
webapp["Django webapp\n(local dev)"]
phpApp["PHP webapp\n(shared hosting)"]
browser(["browser"])
pg --> sqliteExport --> sqliteDB --> phpApp
pg --> webapp
phpApp --> browser
webapp --> browser
end
web -->|"/toate-posturile/?pg_page=N"| fetchIdx
web -->|"/joburi/{slug}/"| fetchDetail
web -->|"wp-content/uploads/"| download
htmlCache --> parseScript
anunturiCSV & calendarCSV --> importCSVs
dlFiles --> extractCmd
pg -->|"body_markdown\n+ attachment_text"| llmSchema
llmSchema -->|"schema_json"| pg
pg --> webapp
python pipeline.py# Scrape, parse, import, LLM enrich, export SQLite — then push to shared host
python pipeline.py && ./deploy-php.sh user@host ~/posturi.gov2.ro
# Quick refresh: skip slow steps that haven't changed
python pipeline.py --skip download,extract,infer,schema && ./deploy-php.sh user@host ~/posturi.gov2.ropython pipeline.py --steps fetch-index,parse,import # specific steps
python pipeline.py --skip download,infer # skip slow steps
python pipeline.py --steps infer --provider gemini # LLM inference only
python pipeline.py --no-llm --force # re-run dict-only inference
python pipeline.py --continue-on-error # log failures, keep going| Step | Script / command | Output |
|---|---|---|
fetch-index |
fetch-index.py |
data/posturi_gov_ro.csv |
fetch-detail |
fetch-anunturi.py |
data/anunturi/**/*.html |
parse |
parse-anunturi.py |
data/anunturi/anunturi.csv + data/calendar.csv |
download |
download-attachments.py |
data/downloads/ |
import |
manage.py import_csvs |
Postgres jobs_jobposting table (normalises județe — see below) |
extract |
manage.py extract_attachments |
JobPosting.attachment_text |
schema |
llm-schema.py |
JobPosting.schema_json JSONB |
infer |
manage.py infer_postings |
JobPosting.inferred JSONB |
occupations |
normalize-titles.py |
jobs_occupation + JobPosting.occupation |
salary |
estimate-salaries.py |
JobPosting.salary_estimate JSONB |
fetch-detail is a bounded refresh, not a one-shot cache: a cached page whose last
successful retrieval is older than --refresh-hours (default 24) is re-fetched, at most
--max-refresh (default 200) per run. Each page gets a <slug>.meta.json sidecar beside
its .html (url, fetched_at, content_hash, bytes, status), written atomically. A cached page with no sidecar
(pre-FIX-03) counts as due — unknown age is not fresh. Index-cancelled postings are skipped.
Due pages are ranked before the cap applies — open competitions first, newest first — and
a competition whose index expiry is more than --refresh-expired-days (default 30) past
is not refreshed at all. The step exits 1 when more than --max-failure-share (default
0.5) of at least five attempted fetches failed, which blocks the deploy like any failed step.
When a refreshed page's Detail Content Hash differs from the stored one, import
requeues what was derived from it: attachment_text is cleared (the extract step
refills it from the cache) and inferred gets a stale_revision marker that infer
selects and replaces — the previous values stay visible until then. Schema extraction
needs no marker (--resume compares revisions itself); occupation and salary recompute
every row each run. A legacy row's first hash is "unknown", not a change. Calendar
events of a posting whose re-parsed page has no schedule are removed.
schema runs before infer: infer_postings reads schema_json for the
announced salary, so the other order leaves inferred.salary_min one run stale.
--force re-processes already-done rows for import, extract, infer, schema and
occupations, and passes --clear to salary (use it after rebuilding the grid).
--limit N restricts infer, schema and occupations to N items (useful for testing).
--provider gemini|openai|anthropic|deepseek sets the LLM used by the infer, schema and
occupations steps (default: gemini).
--no-llm skips the LLM portion of infer and drops occupations to --link-only
(postings are still linked to titles already in the dictionary, which needs no calls).
The schema step is always LLM-driven; use --skip schema to omit it.
A full run is thousands of LLM calls. --compare writes only to
jobs_jobpostingschemavariant and never touches jobs_jobposting.schema_json,
so a sample changes nothing the site serves and the nightly keeps using
whatever LLM_PROMPT_VERSION says.
python llm-schema.py --prompt-version v4 --compare --model-filter deepseek \
--active-only --limit 200 --workers 4
python ops/check-v4-sample.py--limit takes the newest postings first, so a sample is reproducible and
covers what the site is actually serving. --model-filter matters: without it
--compare fans out to every model marked enabled in models_config.json.
check-v4-sample.py reports the deadline-extraction rate, how far expires_at
overstates it, how often v4 recovers a deadline the scraper missed and whether
the two ever disagree, then exits non-zero if the sample is too small, the
deadline rate too low, or the output hit the token budget — truncation is
silent, so it has to be checked rather than noticed. It also flags calendar rows
whose label says "rezultate" but whose stage does not, which is what an enum gap
looks like from the outside.
Romanian public-sector postings essentially never state a salary — 44 of 9,757 do in the body, 5 of 917 attachments — so pay is derived from the 2026 draft salary law rather than scraped.
docs/salarii/…xlsx --build-salary-grid.py--> data/salarii/grila-<versiune>.csv
job title --normalize-titles.py---> jobs_occupation (COR + grid selector)
employer --uat.py----------------> population band
|
v
salary_grid.estimate() -> JobPosting.salary_estimate
The LLM emits a grid selector, never a number. The law is an unadopted draft with more than one public variant, so re-costing under a new one must be a script run, not a re-extraction of 9,600 postings:
python build-salary-grid.py --version-id 2026-08-20 # rebuild the registry
python pipeline.py --steps occupations,salary,export-sqlite --forcedata/salarii/ and data/ocupatii/ are git-tracked on purpose: every figure the
site shows must stay traceable to a sheet and row, including under an older
variant. python build-salary-grid.py --report prints coverage and the rows
flagged for manual review.
Every estimate is gross, at gradation 0, and carries its legal status. Net pay, sporuri and the 20% cap are deliberately out of scope — see the activity log.
The source badge has two shapes — Timiş before the 2026-07 redesign,
TIMIŞOARA, Timiș after — and storing both verbatim once gave the database 261
Judet rows for a country with 42 counties, which quietly broke the județ
filter. webapp/apps/jobs/judete.py is now the single place that interprets a
raw county string:
normalize_judet(raw)→(county, locality), folding the Turkish cedillaş/ţonto Romanianș/țand splitting the city out intoJobPosting.locality.import_csvsapplies it on the way in, so new data cannot fragment.manage.py normalize_judete [--dry-run]repairs data imported before the fix. Idempotent.
When a county cannot be matched, the raw value is kept in
JobPosting.judet_raw and the posting gets no county — so it answers to no
county filter. Three things surface that:
import_csvsprints a warning listing the unmatched values.judet_sanity_warnings()runs on every import (honours--strict) and fires if theJudettable exceeds 42 rows or any posting is unresolved.- Django admin → Anunțuri → filter Județ — rezolvare → Nerecunoscut.
The fix is normally one line in judete.ALIASES.
Compare multiple LLM providers and prompt versions on the same postings without overwriting production results:
# Run all enabled models (respects enable/disable flags in models_config.json)
python llm-schema.py --compare --limit 10
# Test specific models by regex
python llm-schema.py --compare --model-filter "gemini-.*" --limit 5
python llm-schema.py --compare --model-filter "gpt-.*" --limit 5
# Test different prompt versions (when multiple versions exist in config)
python llm-schema.py --compare --prompt-version v2 --limit 10
# Combine: test GPT models with prompt v2
python llm-schema.py --model-filter "gpt-.*" --prompt-version v2 --limit 3Each variant is stored in JobPostingSchemaVariant with:
- Provider, model, prompt version
- Token counts (input/output)
- Cost (USD) calculated from config pricing
- Latency (ms)
View results:
- Django admin:
Admin → Schema LLM variants(filter by provider/model/date) - Job detail page: click "Dev → Comparație LLM" to see all variants for a posting side-by-side
Configuration (models_config.json):
- Models: enable/disable flag per model, pricing (including
cache_input_cost_per_millionfor cache-hit billing) - Prompts: versioned prompts (v1, v2, etc.) centralized in config
get_enabled_models()respects"enabled": true/falseflags
The v2 prompt extracts a flat superset of Schema.org JobPosting properties — keys named after JobPosting properties where they exist (responsibilities, educationRequirements, experienceRequirements, qualifications, skills, baseSalary, jobBenefits, workHours, jobLocation), plus three RO-government-specific custom keys (application_docs, application_fee, application_contact). baseSalary, application_fee, and application_contact are structured dicts; the rest are markdown strings or null.
Pydantic models in schema_models.py are the single source of truth and feed each provider's native structured-output API:
- OpenAI:
response_format={"type":"json_schema","strict":True,...} - Gemini:
response_schema=JobPostingExtraction - Anthropic: tool-use with
input_schema - DeepSeek:
response_format={"type":"json_object"}(loose) + Pydantic post-validation
A cacheable system prefix (instructions + 2 few-shot examples) is sent on every call so providers can hit their prompt cache — measured ~94–99% input-cache hit rate by the 2nd call on OpenAI/DeepSeek. Gemini is the exception: its implicit cache is best-effort and hit only 1 of 6 sequential v3 calls (2026-09-08), so do not budget for it there. See docs/llm-extraction-round.md.
Provider, model and prompt version resolve the same way in llm-schema.py, pipeline.py and quality_check.py — CLI flag > environment variable > models_config.json "defaults", implemented once in llm_config.py:
# .env
LLM_PROVIDER=gemini # gemini | openai | anthropic | deepseek
LLM_MODEL=gemini-2.5-flash # ignored if it is not a model of the selected provider
LLM_PROMPT_VERSION=v4 # v1 | v2 | v3 | v4 (occupation_v1 is the occupation step's prompt)A backfill is ~9,600 calls per model, so the runner is built for that:
# concurrent, restartable, retrying
python llm-schema.py --prompt-version v3 --workers 8 --resume--workers N(default 4) — concurrent LLM calls; database writes stay single-threaded. Measured 5.13 s/post at 4 workers against 15–20 s sequential.--resume— re-pays only for provably stale extractions: no schema yet; a recorded source revision that differs from the posting's current content hash (both known); or recorded provenance with a different provider/model/prompt-version. Unknown is not stale — rows extracted before migration 0014 (no provenance) and rows whose page has not been hashed yet are skipped, so the unattended run never turns into a backfill. A confirmed source change still re-extracts a legacy row: import stamps its pre-change hash.--upgrade-legacy— with--resume, also re-extracts production rows that have no provenance. This is the reviewed, paid FIX-05-RUN backfill (python llm-schema.py --active-only --resume --upgrade-legacy --prompt-version v4 --workers 8); never put it in the cron.--max-attempts N(default 4) — transient errors (429/5xx/timeout) back off exponentially; invalid output gets one repair attempt that shows the model its own validation error.
Exit status: 0 healthy (an empty selection is a healthy no-op); 1 systemic failures (HTTP errors, timeouts, anything unclassified) exceed --max-failure-share (default 0.5) of the attempted postings — or, once a run attempted at least 20 postings, all failures together do, or at least 20 selected postings had no usable content. Per-posting content failures (OutputTruncated, validation/parse errors) are reselected every run, so a small residue of them on a quiet run (a weekend has no new postings) never blocks the deploy, while a broken prompt or model change on a normal run still does; 2 a provider-wide fatal error (HTTP 401/402/403), after which no new calls are scheduled. Each model's run summary is appended to data/pipeline-runs.jsonl as kind: "llm-schema".
Every run also checks that quoted evidence/verbatim values actually appear in the posting (grounding.py) and reports unsupported quotes — a hallucination check with no extra LLM call.
Before sending to the LLM, boilerplate.py::strip_hg_1336() removes generic eligibility lines from HG 1.336/2022 / Codul muncii / OUG 57/2019 art. 542 (cetățenia română, capacitate de muncă, condamnări, pedepse complementare, condițiile generice de studii/vechime etc.). These appear nearly verbatim on every posting and otherwise drown the qualifications field with legal citation. The art. 35 dosar list survives intact (it's the application_docs content). Measured 13–25% input-length reduction on sampled postings; the bigger win is qualifications becoming role-specific signal (76–295 chars) instead of ~2300 chars of legal boilerplate. Toggle off with python llm-schema.py --no-strip.
quality_check.py samples 5–10 diverse postings and runs all four pipeline layers through automated checks, then writes data/quality_report.json and a console summary table.
# Fast pass — no API calls (CSV fields, attachment extraction, dict-only inference)
webapp/.venv/bin/python3 quality_check.py --no-llm
# Full pass — includes LLM infer fallback + schema.org generation
webapp/.venv/bin/python3 quality_check.py --provider anthropic
# Check specific postings by slug
webapp/.venv/bin/python3 quality_check.py --slugs 2a66f376.doc,67438cc9.docxUse the webapp venv because it has docx2txt, python-docx, and pypdf. PDF OCR fallback requires system tools: brew install poppler tesseract tesseract-lang (poppler provides pdftoppm; tesseract-lang installs ron for Romanian). After running, invoke /quality-review in Claude Code for a narrative assessment with root-cause analysis and recommended fixes.
data/posturi_gov_ro.csv — job index, one row per listing, keyed by URL:
| Field | Description |
|---|---|
pozitie |
Job title |
url |
Listing URL (primary key) |
angajator |
Employer |
detalii |
Details (comma-separated tags) |
publicat_in |
Publication date |
expira_in |
Expiry date |
judet |
County |
url_judet |
County filter URL |
tip |
Listing type |
updates |
Semicolon-separated log of field changes with dates |
data/anunturi/anunturi.csv — one row per cached announcement:
| Field | Description |
|---|---|
Job Title |
Position title |
Employer |
Hiring organisation |
Location |
County / locality |
Job Level |
Funcții de execuție / Funcții de conducere |
Job Type |
Permanent / Temporar |
Employer Category |
Angajator type (Primării, Instituții locale, Guvern și ministere, etc.) |
Categorie |
Funcție contractuală / Funcție publică |
Announcement URL |
Link to attached document or posting page |
Main Body Markdown |
Full announcement text converted to markdown |
Other Links |
Comma-separated attachment URLs |
Nr Posturi |
Number of vacancies |
Contact Telefon |
Phone number extracted from body text |
Contact Email |
Email address extracted from body text |
Contact Persoana |
Contact person name |
Data Limita Depunere |
Application deadline (DD.MM.YYYY, ora HH:MM) |
Data Proba Scrisa |
Written test date |
Data Interviu |
Interview date |
Data Rezultate Finale |
Final results date |
Status |
The detail page's own marker: Live / Anulat; blank = unknown |
Detail Fetched At |
When the detail page was last successfully fetched (from the sidecar); blank on legacy rows = unknown |
Detail Content Hash |
Hash of the fetched page, the revision llm-schema.py --resume compares against; blank on legacy rows = unknown |
data/calendar.csv — flat competition timeline table, one row per event:
| Field | Description |
|---|---|
url |
Announcement URL |
eveniment |
Event description (e.g. "Depunerea dosarelor", "Proba scrisă") |
data |
Date in DD.MM.YYYY format |
ora |
Time in HH:MM format (empty if not specified) |
data/ is gitignored.
pip install -r requirements.txtFor LLM scripts and the webapp, copy .env.example to .env and fill in your API keys.
dox2md.py also requires system packages:
brew install libreoffice pandoc tesseractpython3 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/python webapp/manage.py migrate
.venv/bin/python webapp/manage.py runserverA lightweight PHP frontend that runs on commodity shared hosting (cPanel). Reads from a read-only SQLite database — no Python, no web server config, just PHP + SQLite.
The PHP app has its own fixture-based suites under webapp-php/tests/ — no
network, no production data, no paid APIs. Each suite builds a deterministic
SQLite fixture (tests/fixtures/build_db.php) and, where it needs rendered
pages, serves the real app on the built-in PHP router with a fixed clock.
php tests/sanitize_test.php # FIX-01: Markdown sanitizer payload battery
php tests/deadline_test.php # FIX-02: one deadline across list/detail/feeds
php tests/request_test.php # FIX-07: request validation, routes, feeds, filters
php tests/compat_test.php # REV-01: new code against an older export's schema
npx playwright test --config webapp-php/tests/browser/playwright.config.js
# FIX-07: browser checks at 320/375/1280These run locally. .github/workflows/ci.yml exists (lint + PHP suites on PHP 8.2,
Playwright with pinned Chromium, alongside the Django/PostgreSQL suite) but GitHub
Actions is not activated — CI was postponed by decision. The
fixture query decoder is webapp-php/query.php — every request passes through
validated_query() at the front controller, and ?q[]=medic (a 500 before
FIX-07) is a controlled 400.
# Everything: rebuild the SQLite from Postgres, push code and data
./deploy-php.sh user@host '~/posturi.gov2.ro'
# Just the PHP tree — after a template or stylesheet change
./deploy-php.sh --code-only
# Just the database, already built by pipeline.py — what the cron runs
./deploy-php.sh --data-only --no-export
# See what would move, change nothing
./deploy-php.sh --dry-run
# Push nothing; compare local vs served code/data markers (needs SITE_URL)
./deploy-php.sh --verify # exit 1 when a code or data release is pendingSet DEPLOY_HOST, DEPLOY_PATH and SITE_URL in .env and the arguments become
optional. Quote a leading ~: unquoted it expands against the local home, and the
script refuses the result rather than rsyncing to a path the remote host has never
heard of.
The deploy script:
- Aborts if
webapp-php/static/app.cssis missing (see Stylesheet below), or ifDEPLOY_PATHlooks like a home directory — it runsrsync --delete - Runs
export-to-sqlite.py --active-only— pulls active postings (expires_at >= today) from PostgreSQL intowebapp-php/posturi.sqlite, unless--no-export - Refuses to ship a database that fails
integrity_checkor has no postings - Code: rsyncs
webapp-php/with--delete, minusassets/(Tailwind source),router.phpandtests/(dev only) and*.sqlite*— excluding the database also protects it from the deletion pass - Data: rsyncs
posturi.sqlitealone, no--deleteand no--inplace, so rsync's write-temp-then-rename swaps it atomically and a request mid-transfer still sees the whole previous database - Code deploy only: stamps the current
git rev-parse --short HEADintowebapp-php/static/code-version.txt(gitignored) before the rsync, so the host names the checkout it runs - Checks
SITE_URLreturns HTTP 200, then verifies the served markers on/versiuni.json—codeafter a code push,data.built_atafter a data push — with up to 5 attempts 5 s apart, since the shared host may serve the old file for a moment. A mismatch fails the deploy. All of step 7 is skipped whenSITE_URLis unset
/versiuni.json (webapp-php/feeds/versiuni.json.php, Cache-Control: no-store) reports
the deployed code marker and the data provenance (built_at, run_id, git_sha,
postings, active_only, detail_fetched_at_max, detail_fetched_rows,
index_checked_at) as two separate facts; source_host is not exposed. --verify compares
the same two markers against this checkout and the local posturi.sqlite.
Order matters across a schema change. Code and data deploy separately, so new PHP
code can land on an export written by older pipeline code. db() covers the known gap:
when job_postings lacks application_status (exports before FIX-05), it shadows the
table with a TEMP view that computes the status exactly as the export does
(shim_legacy_export() in webapp-php/db.php, tested by tests/compat_test.php). Any
new column the templates start reading needs the same treatment or a VPS-first
rollout: upgrade the VPS checkout, migrate, export and --data-only first, then
--code-only.
The full archive stays in PostgreSQL; the deployed SQLite only contains currently active postings.
The export builds beside its target and only replaces it once it opens, passes
integrity_check, clears --min-rows (default 100) and has not lost half its rows
against the file it would replace. --force overrides the row floors. This is the guard
that was missing when the live site served three active postings for five weeks.
Each export carries a build_meta row — when it was built, from which commit and host,
and what it holds. FIX-06 added the provenance fields: run_id (the pipeline run, from
POSTURI_RUN_ID), index_checked_at / index_scan_pages / index_scan_complete (from
data/index-scan.json), and detail_fetched_at_max / detail_fetched_rows (newest
Detail Fetched At and how many rows have one); all are empty/null on legacy builds =
unknown, never synthesized from build time. The site reads it for the provenance block on /despre/ and for the
tooltip behind the "actualizat" stamps; build_meta() in helpers.php returns null on an
older export, and every consumer renders without it. Note the two timestamps mean
different things: the visible stamp is MAX(last_seen_at), when the source was last
scraped, while built_at is when the file was generated. source_host is recorded for
the deploy to verify against, never rendered.
job_postings.application_status is the one vocabulary every surface filters on:
confirmed_open (deadline from the concurs or the announcement, still ahead), unconfirmed
(only the listing's expiry date, still ahead), closed (any date in the past) and unknown
(no date). Cancelled postings are excluded from the export rather than given a status.
The export also carries llm_costs — one row per Bucharest day, provider, model and
prompt version, summed from jobs_jobpostingschemavariant — which feeds the "Costuri
inferență" section of /statistici. It is not cut down by --active-only, and it covers
the schema extraction step only: infer and occupations record no tokens or cost. A
variant is upserted, so a re-extracted posting's cost moves to its newest day. An export
that predates the table makes the section read "Date indisponibile" instead of failing.
The pipeline runs unattended on a VPS twice a day (13:15 and 18:33 Europe/Bucharest) and
pushes a fresh database to the shared host. The repo ships a systemd timer pinned to
Europe/Bucharest (ops/systemd/posturi-pipeline.timer, needs systemd >= 240); the live
box still runs a user crontab until OPS-01 is completed. Code deploys stay manual from the development
machine, which is why the two rsyncs above are separate — a cron that pushed the whole
directory would revert templates from the VPS's older checkout.
Mac (dev) ──git push──> GitHub ──git pull (manual)──> VPS
│ │
│ ./deploy-php.sh --code-only ops/run-pipeline.sh (timer / cron)
└──────────────> shared host (PHP + posturi.sqlite) <┘
| Path | What it is |
|---|---|
ops/run-pipeline.sh |
the unattended entry point: flock, pipeline, export check, data deploy, healthcheck ping |
ops/check-export.py |
the deploy gate — asserts the SQLite about to ship is fit to ship |
ops/systemd/posturi-pipeline.{service,timer} |
the two daily slots (not what the live box runs yet; the units hard-code User=posturi and /srv/posturi) |
ops/env.sh |
.env reader shared by the shell scripts (it is never sourced — it holds API keys) |
ops/logrotate.posturi |
rotation for logs/pipeline.log, installed into /etc/logrotate.d/ |
docs/deploy-vps.md |
provisioning runbook, operating commands, failure table |
Each run leaves three things behind:
| Where | What |
|---|---|
data/pipeline-runs.jsonl |
one kind: "run" record (step timings, exit codes, flags) and one kind: "export-check" record (33 data metrics, every assertion), joined by run_id; plus kind: "llm-balance" (the pre-flight balance and its outcome) |
logs/pipeline.log |
the steps' own output, runs delimited by === posturi pipeline run <id> … === |
HEALTHCHECK_URL |
/start, /fail, success — the only signal that can report a run which never happened |
By default a failed pipeline step blocks the deploy: ops/run-pipeline.sh pings /fail, exits
with the pipeline's status and the host keeps serving the previous database. Set
POSTURI_ALLOW_DEGRADED_DEPLOY=1 to deploy anyway when the export check passes (the run is
recorded as degraded).
ops/check-llm-balance.py runs first, before anything is fetched or paid for: it reads
the DeepSeek balance (GET /user/balance, free) and exits 69 below LLM_BALANCE_MIN
(default 0.50 USD, or when the account reports is_available: false), which
ops/run-pipeline.sh turns into a /fail ping and an immediate stop. Below
LLM_BALANCE_WARN (default 3.00) it only warns. A check that cannot reach the API,
or a provider other than DeepSeek, prints a line and exits 0 — it never blocks a run.
POSTURI_SKIP_BALANCE_CHECK=1 bypasses it; --need USD is for a manual backfill of known
cost. Exit codes of ops/run-pipeline.sh: 0 fine, 65 export failed its hard checks,
69 balance too low, 75 another run holds the lock, 78 unapplied migrations.
ops/check-export.py runs between the export and the deploy. Hard checks abort
before the rsync, so a corrupt export cannot reach the live site: integrity_check,
broken foreign keys, duplicate URLs, a short FTS index, anything other than exactly 42
județ rows, a collapse below half the previous run, or a build_meta.built_at older
than six hours (which means the export step failed and the deploy is about to ship an
old file). Soft checks — coverage regressions, an altele spike, one profession
family or skill tag taking an implausible share, anomaly-flag step changes, mojibake —
are recorded and printed but do not block; --strict makes them fatal.
Every check is tied to a regression that actually happened here, and the two
*_sanity_warnings() functions from import_csvs.py are called rather than restated,
so "too many counties" has one definition in the tree.
.venv/bin/python ops/check-export.py --no-log # read the live numbers
.venv/bin/python ops/check-export.py --db /tmp/copy.sqlite --no-log--no-log keeps an interactive read from becoming the baseline the next real run
compares against. For the narrative pass over the run log — missed slots, recurring
step failures, metrics sliding while every individual check still passes — run
/pipeline-check in Claude Code. See docs/pipeline-quality-checks.md for the
full watch-list and what is still unbuilt.
Full setup — packages, seeding Postgres and the scrape cache from the Mac, SSH keys, installing the timer — is in docs/deploy-vps.md.
The app has no build step on the host — the compiled CSS is committed. Rebuild it whenever you touch a template's classes:
npm install # once
npm run css # webapp-php/assets/app.css -> webapp-php/static/app.css (minified)
npm run css:watch # while editing templatesFonts (static/fonts/*.woff2) and htmx (static/htmx.min.js) are self-hosted; nothing
is fetched from a CDN at runtime.
The whole palette is one :root block of CSS custom properties in
webapp-php/assets/app.css. Tailwind's theme resolves every colour, radius and font
utility to one of those variables, so re-declaring them under [data-skin="<id>"]
restyles the entire site without touching a single utility class. That is all a skin is.
Three ship today, chosen with the picker in the footer (stored in localStorage, applied
before first paint so there is no flash):
| id | what it is |
|---|---|
hartie |
the built-in default — parchment, Fraunces, DM Sans. No file; it is :root. |
govuk |
GOV.UK Design System — white, square, Arial, yellow focus, black masthead. |
posturi |
the official posturi.gov.ro — navy + gold, Manrope, rounded, lifted cards. |
Adding one: copy webapp-php/static/skins/_template.css to <name>.css and it
appears in the picker on the next request — no registry, no build step. The id is the
filename; inc/skins.php discovers it and reads the display name from the @skin comment.
Files starting with _ are skipped.
Two rules the template explains at length:
- Scope every rule under
[data-skin="<id>"]. All skin files load on every page, so an unscoped rule leaks into the others.@font-faceis the exception — it declares a family rather than applying one, and is how a skin ships its own typeface. - Colours are space-separated RGB channels (
245 240 232), not hex. That is the only form Tailwind's alpha modifiers compose with; a hex value silently breaks everybg-surface/70on the page.
Skins live in static/ rather than assets/ because they need no build — Tailwind would
strip them, since nothing in the PHP references their selectors — and because deploy-php.sh
excludes assets/.
Validate before committing:
php webapp-php/assets/check-skins.phpIt reads the token contract out of app.css and flags the four failures that are silent in
a browser: an unscoped rule, a token name that does not exist, a colour written as hex, and
any text/background pair under WCAG AA.
db.php honours a POSTURI_DB environment variable, so a test or a preview can
point at a different SQLite file without touching the deployed one:
POSTURI_DB=/tmp/other.sqlite php -S localhost:8000 -t webapp-php webapp-php/router.phpUnset in production.
php -S localhost:8000 -t webapp-php webapp-php/router.phprouter.php exists only for PHP's built-in server — it serves static files and mirrors
the .htaccess deny rules. Apache never loads it.
| Path | Purpose |
|---|---|
index.php |
Front controller — routes /, /job/123-slug/, /angajatori/, /statistici/, /despre/, /robots.txt, /sitemap.xml |
db.php |
PDO singleton for posturi.sqlite |
helpers.php |
Markdown rendering, date formatting, filter builder, facet queries, FTS query builder, active-filter chips, URL helpers |
pages/ |
List, detail, employer list, employer detail, stats, about |
feeds/ |
Atom, JSON API, iCal endpoints — follow the active filters; ?employer=<slug> scopes a feed to one angajator (linked from each employer profile) |
partials/ |
Result list and facet sidebar (both HTMX out-of-band swappable) |
inc/ |
Header/footer HTML, <head> metadata (canonical, OpenGraph, JSON-LD hook) |
assets/ |
Tailwind source — not deployed |
static/ |
Compiled app.css, self-hosted fonts, htmx |
router.php |
Dev-server router — not deployed |
.htaccess |
URL rewriting, asset caching, blocks direct access to *.sqlite* |