Skip to content

About

posturi.gov.ro, da mai drăgu'

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

207 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

posturi.gov.ro scraper

→ posturi.gov2.ro [WIP]

Alternative browser / explorer for posturi.gov.ro. Scrapes the Romanian government job listings portal — and tracks changes over time. Pipeline: index → cache announcement pages → extract structured data → LLM-extracted structured display sections stored in Postgres.

Derivative work: mariuscomper.uk/posturi-publice

Project review and next work

The October 3 project audit records the code, documentation and live-site findings. Use the consolidated backlog for current priorities and the remediation specifications for bounded coding-agent assignments. Historical tracking is preserved in the activity log and the pre-consolidation backlog.

The agreed product direction captures the user-facing feature scope and recommended sequence toward an application companion.

Pipeline

flowchart LR
    web(["posturi.gov.ro"])

    subgraph scrape["① Scrape"]
        direction TB
        fetchIdx["fetch-index.py"]
        indexCSV[/"posturi_gov_ro.csv"/]
        fetchDetail["fetch-anunturi.py"]
        htmlCache[/"anunturi/\n**∕*.html"/]
        download["download-attachments.py"]
        dlFiles[/"downloads/\n*.docx *.pdf"/]

        fetchIdx --> indexCSV --> fetchDetail --> htmlCache
        htmlCache --> download --> dlFiles
    end

    subgraph parse["② Parse"]
        direction TB
        parseScript["parse-anunturi.py"]
        anunturiCSV[/"anunturi.csv"/]
        calendarCSV[/"calendar.csv"/]
        parseScript --> anunturiCSV & calendarCSV
    end

    subgraph db_layer["③ Import → Postgres"]
        direction TB
        importCSVs["import_csvs"]
        extractCmd["extract_attachments"]
        inferCmd["infer_postings"]
        pg[("jobs_jobposting\nbody_markdown\nattachment_text\ninferred JSONB")]

        importCSVs --> pg
        extractCmd -->|"attachment_text"| pg
        pg --> inferCmd -->|"inferred"| pg
    end

    subgraph enrich["④ Enrich"]
        llmSchema["llm-schema.py"]
    end

    subgraph serve["⑤ Serve"]
        direction TB
        pg[/"PostgreSQL"/]
        sqliteExport["export-to-sqlite.py\n--active-only"]
        sqliteDB[/"posturi.sqlite\n(active only)"/]
        webapp["Django webapp\n(local dev)"]
        phpApp["PHP webapp\n(shared hosting)"]
        browser(["browser"])
        
        pg --> sqliteExport --> sqliteDB --> phpApp
        pg --> webapp
        phpApp --> browser
        webapp --> browser
    end

    web -->|"/toate-posturile/?pg_page=N"| fetchIdx
    web -->|"/joburi/{slug}/"| fetchDetail
    web -->|"wp-content/uploads/"| download

    htmlCache --> parseScript
    anunturiCSV & calendarCSV --> importCSVs
    dlFiles --> extractCmd

    pg -->|"body_markdown\n+ attachment_text"| llmSchema
    llmSchema -->|"schema_json"| pg
    pg --> webapp
Loading

Quick start — run everything

python pipeline.py

Full update → deploy

# Scrape, parse, import, LLM enrich, export SQLite — then push to shared host
python pipeline.py && ./deploy-php.sh user@host ~/posturi.gov2.ro

# Quick refresh: skip slow steps that haven't changed
python pipeline.py --skip download,extract,infer,schema && ./deploy-php.sh user@host ~/posturi.gov2.ro

Selective runs

python pipeline.py --steps fetch-index,parse,import   # specific steps
python pipeline.py --skip download,infer               # skip slow steps
python pipeline.py --steps infer --provider gemini     # LLM inference only
python pipeline.py --no-llm --force                    # re-run dict-only inference
python pipeline.py --continue-on-error                 # log failures, keep going

Steps

Step Script / command Output
fetch-index fetch-index.py data/posturi_gov_ro.csv
fetch-detail fetch-anunturi.py data/anunturi/**/*.html
parse parse-anunturi.py data/anunturi/anunturi.csv + data/calendar.csv
download download-attachments.py data/downloads/
import manage.py import_csvs Postgres jobs_jobposting table (normalises județe — see below)
extract manage.py extract_attachments JobPosting.attachment_text
schema llm-schema.py JobPosting.schema_json JSONB
infer manage.py infer_postings JobPosting.inferred JSONB
occupations normalize-titles.py jobs_occupation + JobPosting.occupation
salary estimate-salaries.py JobPosting.salary_estimate JSONB

fetch-detail is a bounded refresh, not a one-shot cache: a cached page whose last successful retrieval is older than --refresh-hours (default 24) is re-fetched, at most --max-refresh (default 200) per run. Each page gets a <slug>.meta.json sidecar beside its .html (url, fetched_at, content_hash, bytes, status), written atomically. A cached page with no sidecar (pre-FIX-03) counts as due — unknown age is not fresh. Index-cancelled postings are skipped. Due pages are ranked before the cap applies — open competitions first, newest first — and a competition whose index expiry is more than --refresh-expired-days (default 30) past is not refreshed at all. The step exits 1 when more than --max-failure-share (default 0.5) of at least five attempted fetches failed, which blocks the deploy like any failed step.

When a refreshed page's Detail Content Hash differs from the stored one, import requeues what was derived from it: attachment_text is cleared (the extract step refills it from the cache) and inferred gets a stale_revision marker that infer selects and replaces — the previous values stay visible until then. Schema extraction needs no marker (--resume compares revisions itself); occupation and salary recompute every row each run. A legacy row's first hash is "unknown", not a change. Calendar events of a posting whose re-parsed page has no schedule are removed.

schema runs before infer: infer_postings reads schema_json for the announced salary, so the other order leaves inferred.salary_min one run stale.

--force re-processes already-done rows for import, extract, infer, schema and occupations, and passes --clear to salary (use it after rebuilding the grid). --limit N restricts infer, schema and occupations to N items (useful for testing). --provider gemini|openai|anthropic|deepseek sets the LLM used by the infer, schema and occupations steps (default: gemini). --no-llm skips the LLM portion of infer and drops occupations to --link-only (postings are still linked to titles already in the dictionary, which needs no calls). The schema step is always LLM-driven; use --skip schema to omit it.

Sampling a prompt before a corpus-wide run

A full run is thousands of LLM calls. --compare writes only to jobs_jobpostingschemavariant and never touches jobs_jobposting.schema_json, so a sample changes nothing the site serves and the nightly keeps using whatever LLM_PROMPT_VERSION says.

python llm-schema.py --prompt-version v4 --compare --model-filter deepseek \
    --active-only --limit 200 --workers 4
python ops/check-v4-sample.py

--limit takes the newest postings first, so a sample is reproducible and covers what the site is actually serving. --model-filter matters: without it --compare fans out to every model marked enabled in models_config.json.

check-v4-sample.py reports the deadline-extraction rate, how far expires_at overstates it, how often v4 recovers a deadline the scraper missed and whether the two ever disagree, then exits non-zero if the sample is too small, the deadline rate too low, or the output hit the token budget — truncation is silent, so it has to be checked rather than noticed. It also flags calendar rows whose label says "rezultate" but whose stage does not, which is what an enum gap looks like from the outside.

Estimated pay (salarii/, ocupatii/)

Romanian public-sector postings essentially never state a salary — 44 of 9,757 do in the body, 5 of 917 attachments — so pay is derived from the 2026 draft salary law rather than scraped.

docs/salarii/…xlsx  --build-salary-grid.py-->  data/salarii/grila-<versiune>.csv
job title           --normalize-titles.py--->  jobs_occupation (COR + grid selector)
employer            --uat.py---------------->  population band
                                    |
                                    v
                   salary_grid.estimate()  ->  JobPosting.salary_estimate

The LLM emits a grid selector, never a number. The law is an unadopted draft with more than one public variant, so re-costing under a new one must be a script run, not a re-extraction of 9,600 postings:

python build-salary-grid.py --version-id 2026-08-20   # rebuild the registry
python pipeline.py --steps occupations,salary,export-sqlite --force

data/salarii/ and data/ocupatii/ are git-tracked on purpose: every figure the site shows must stay traceable to a sheet and row, including under an older variant. python build-salary-grid.py --report prints coverage and the rows flagged for manual review.

Every estimate is gross, at gradation 0, and carries its legal status. Net pay, sporuri and the 20% cap are deliberately out of scope — see the activity log.

Județe (counties)

The source badge has two shapes — Timiş before the 2026-07 redesign, TIMIŞOARA, Timiș after — and storing both verbatim once gave the database 261 Judet rows for a country with 42 counties, which quietly broke the județ filter. webapp/apps/jobs/judete.py is now the single place that interprets a raw county string:

  • normalize_judet(raw) → (county, locality), folding the Turkish cedilla ş/ţ onto Romanian ș/ț and splitting the city out into JobPosting.locality.
  • import_csvs applies it on the way in, so new data cannot fragment.
  • manage.py normalize_judete [--dry-run] repairs data imported before the fix. Idempotent.

When a county cannot be matched, the raw value is kept in JobPosting.judet_raw and the posting gets no county — so it answers to no county filter. Three things surface that:

  1. import_csvs prints a warning listing the unmatched values.
  2. judet_sanity_warnings() runs on every import (honours --strict) and fires if the Judet table exceeds 42 rows or any posting is unresolved.
  3. Django admin → Anunțuri → filter Județ — rezolvare → Nerecunoscut.

The fix is normally one line in judete.ALIASES.

LLM Provider Comparison

Compare multiple LLM providers and prompt versions on the same postings without overwriting production results:

# Run all enabled models (respects enable/disable flags in models_config.json)
python llm-schema.py --compare --limit 10

# Test specific models by regex
python llm-schema.py --compare --model-filter "gemini-.*" --limit 5
python llm-schema.py --compare --model-filter "gpt-.*" --limit 5

# Test different prompt versions (when multiple versions exist in config)
python llm-schema.py --compare --prompt-version v2 --limit 10

# Combine: test GPT models with prompt v2
python llm-schema.py --model-filter "gpt-.*" --prompt-version v2 --limit 3

Each variant is stored in JobPostingSchemaVariant with:

  • Provider, model, prompt version
  • Token counts (input/output)
  • Cost (USD) calculated from config pricing
  • Latency (ms)

View results:

  • Django admin: Admin → Schema LLM variants (filter by provider/model/date)
  • Job detail page: click "Dev → Comparație LLM" to see all variants for a posting side-by-side

Configuration (models_config.json):

  • Models: enable/disable flag per model, pricing (including cache_input_cost_per_million for cache-hit billing)
  • Prompts: versioned prompts (v1, v2, etc.) centralized in config
  • get_enabled_models() respects "enabled": true/false flags

Prompt v2 (Schema.org-aligned)

The v2 prompt extracts a flat superset of Schema.org JobPosting properties — keys named after JobPosting properties where they exist (responsibilities, educationRequirements, experienceRequirements, qualifications, skills, baseSalary, jobBenefits, workHours, jobLocation), plus three RO-government-specific custom keys (application_docs, application_fee, application_contact). baseSalary, application_fee, and application_contact are structured dicts; the rest are markdown strings or null.

Pydantic models in schema_models.py are the single source of truth and feed each provider's native structured-output API:

  • OpenAI: response_format={"type":"json_schema","strict":True,...}
  • Gemini: response_schema=JobPostingExtraction
  • Anthropic: tool-use with input_schema
  • DeepSeek: response_format={"type":"json_object"} (loose) + Pydantic post-validation

A cacheable system prefix (instructions + 2 few-shot examples) is sent on every call so providers can hit their prompt cache — measured ~94–99% input-cache hit rate by the 2nd call on OpenAI/DeepSeek. Gemini is the exception: its implicit cache is best-effort and hit only 1 of 6 sequential v3 calls (2026-09-08), so do not budget for it there. See docs/llm-extraction-round.md.

Choosing the model, and long runs

Provider, model and prompt version resolve the same way in llm-schema.py, pipeline.py and quality_check.py — CLI flag > environment variable > models_config.json "defaults", implemented once in llm_config.py:

# .env
LLM_PROVIDER=gemini            # gemini | openai | anthropic | deepseek
LLM_MODEL=gemini-2.5-flash     # ignored if it is not a model of the selected provider
LLM_PROMPT_VERSION=v4          # v1 | v2 | v3 | v4 (occupation_v1 is the occupation step's prompt)

A backfill is ~9,600 calls per model, so the runner is built for that:

# concurrent, restartable, retrying
python llm-schema.py --prompt-version v3 --workers 8 --resume
  • --workers N (default 4) — concurrent LLM calls; database writes stay single-threaded. Measured 5.13 s/post at 4 workers against 15–20 s sequential.
  • --resume — re-pays only for provably stale extractions: no schema yet; a recorded source revision that differs from the posting's current content hash (both known); or recorded provenance with a different provider/model/prompt-version. Unknown is not stale — rows extracted before migration 0014 (no provenance) and rows whose page has not been hashed yet are skipped, so the unattended run never turns into a backfill. A confirmed source change still re-extracts a legacy row: import stamps its pre-change hash.
  • --upgrade-legacy — with --resume, also re-extracts production rows that have no provenance. This is the reviewed, paid FIX-05-RUN backfill (python llm-schema.py --active-only --resume --upgrade-legacy --prompt-version v4 --workers 8); never put it in the cron.
  • --max-attempts N (default 4) — transient errors (429/5xx/timeout) back off exponentially; invalid output gets one repair attempt that shows the model its own validation error.

Exit status: 0 healthy (an empty selection is a healthy no-op); 1 systemic failures (HTTP errors, timeouts, anything unclassified) exceed --max-failure-share (default 0.5) of the attempted postings — or, once a run attempted at least 20 postings, all failures together do, or at least 20 selected postings had no usable content. Per-posting content failures (OutputTruncated, validation/parse errors) are reselected every run, so a small residue of them on a quiet run (a weekend has no new postings) never blocks the deploy, while a broken prompt or model change on a normal run still does; 2 a provider-wide fatal error (HTTP 401/402/403), after which no new calls are scheduled. Each model's run summary is appended to data/pipeline-runs.jsonl as kind: "llm-schema".

Every run also checks that quoted evidence/verbatim values actually appear in the posting (grounding.py) and reports unsupported quotes — a hallucination check with no extra LLM call.

Boilerplate stripping

Before sending to the LLM, boilerplate.py::strip_hg_1336() removes generic eligibility lines from HG 1.336/2022 / Codul muncii / OUG 57/2019 art. 542 (cetățenia română, capacitate de muncă, condamnări, pedepse complementare, condițiile generice de studii/vechime etc.). These appear nearly verbatim on every posting and otherwise drown the qualifications field with legal citation. The art. 35 dosar list survives intact (it's the application_docs content). Measured 13–25% input-length reduction on sampled postings; the bigger win is qualifications becoming role-specific signal (76–295 chars) instead of ~2300 chars of legal boilerplate. Toggle off with python llm-schema.py --no-strip.

Data quality

quality_check.py samples 5–10 diverse postings and runs all four pipeline layers through automated checks, then writes data/quality_report.json and a console summary table.

# Fast pass — no API calls (CSV fields, attachment extraction, dict-only inference)
webapp/.venv/bin/python3 quality_check.py --no-llm

# Full pass — includes LLM infer fallback + schema.org generation
webapp/.venv/bin/python3 quality_check.py --provider anthropic

# Check specific postings by slug
webapp/.venv/bin/python3 quality_check.py --slugs 2a66f376.doc,67438cc9.docx

Use the webapp venv because it has docx2txt, python-docx, and pypdf. PDF OCR fallback requires system tools: brew install poppler tesseract tesseract-lang (poppler provides pdftoppm; tesseract-lang installs ron for Romanian). After running, invoke /quality-review in Claude Code for a narrative assessment with root-cause analysis and recommended fixes.

Data

data/posturi_gov_ro.csv — job index, one row per listing, keyed by URL:

Field Description
pozitie Job title
url Listing URL (primary key)
angajator Employer
detalii Details (comma-separated tags)
publicat_in Publication date
expira_in Expiry date
judet County
url_judet County filter URL
tip Listing type
updates Semicolon-separated log of field changes with dates

data/anunturi/anunturi.csv — one row per cached announcement:

Field Description
Job Title Position title
Employer Hiring organisation
Location County / locality
Job Level Funcții de execuție / Funcții de conducere
Job Type Permanent / Temporar
Employer Category Angajator type (Primării, Instituții locale, Guvern și ministere, etc.)
Categorie Funcție contractuală / Funcție publică
Announcement URL Link to attached document or posting page
Main Body Markdown Full announcement text converted to markdown
Other Links Comma-separated attachment URLs
Nr Posturi Number of vacancies
Contact Telefon Phone number extracted from body text
Contact Email Email address extracted from body text
Contact Persoana Contact person name
Data Limita Depunere Application deadline (DD.MM.YYYY, ora HH:MM)
Data Proba Scrisa Written test date
Data Interviu Interview date
Data Rezultate Finale Final results date
Status The detail page's own marker: Live / Anulat; blank = unknown
Detail Fetched At When the detail page was last successfully fetched (from the sidecar); blank on legacy rows = unknown
Detail Content Hash Hash of the fetched page, the revision llm-schema.py --resume compares against; blank on legacy rows = unknown

data/calendar.csv — flat competition timeline table, one row per event:

Field Description
url Announcement URL
eveniment Event description (e.g. "Depunerea dosarelor", "Proba scrisă")
data Date in DD.MM.YYYY format
ora Time in HH:MM format (empty if not specified)

data/ is gitignored.

Setup

pip install -r requirements.txt

For LLM scripts and the webapp, copy .env.example to .env and fill in your API keys.

dox2md.py also requires system packages:

brew install libreoffice pandoc tesseract

Webapp setup

python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/python webapp/manage.py migrate
.venv/bin/python webapp/manage.py runserver

PHP webapp (shared hosting)

A lightweight PHP frontend that runs on commodity shared hosting (cPanel). Reads from a read-only SQLite database — no Python, no web server config, just PHP + SQLite.

Tests

The PHP app has its own fixture-based suites under webapp-php/tests/ — no network, no production data, no paid APIs. Each suite builds a deterministic SQLite fixture (tests/fixtures/build_db.php) and, where it needs rendered pages, serves the real app on the built-in PHP router with a fixed clock.

php tests/sanitize_test.php     # FIX-01: Markdown sanitizer payload battery
php tests/deadline_test.php     # FIX-02: one deadline across list/detail/feeds
php tests/request_test.php      # FIX-07: request validation, routes, feeds, filters
php tests/compat_test.php       # REV-01: new code against an older export's schema

npx playwright test --config webapp-php/tests/browser/playwright.config.js
                                # FIX-07: browser checks at 320/375/1280

These run locally. .github/workflows/ci.yml exists (lint + PHP suites on PHP 8.2, Playwright with pinned Chromium, alongside the Django/PostgreSQL suite) but GitHub Actions is not activated — CI was postponed by decision. The fixture query decoder is webapp-php/query.php — every request passes through validated_query() at the front controller, and ?q[]=medic (a 500 before FIX-07) is a controlled 400.

Export & deploy

# Everything: rebuild the SQLite from Postgres, push code and data
./deploy-php.sh user@host '~/posturi.gov2.ro'

# Just the PHP tree — after a template or stylesheet change
./deploy-php.sh --code-only

# Just the database, already built by pipeline.py — what the cron runs
./deploy-php.sh --data-only --no-export

# See what would move, change nothing
./deploy-php.sh --dry-run

# Push nothing; compare local vs served code/data markers (needs SITE_URL)
./deploy-php.sh --verify        # exit 1 when a code or data release is pending

Set DEPLOY_HOST, DEPLOY_PATH and SITE_URL in .env and the arguments become optional. Quote a leading ~: unquoted it expands against the local home, and the script refuses the result rather than rsyncing to a path the remote host has never heard of.

The deploy script:

  1. Aborts if webapp-php/static/app.css is missing (see Stylesheet below), or if DEPLOY_PATH looks like a home directory — it runs rsync --delete
  2. Runs export-to-sqlite.py --active-only — pulls active postings (expires_at >= today) from PostgreSQL into webapp-php/posturi.sqlite, unless --no-export
  3. Refuses to ship a database that fails integrity_check or has no postings
  4. Code: rsyncs webapp-php/ with --delete, minus assets/ (Tailwind source), router.php and tests/ (dev only) and *.sqlite* — excluding the database also protects it from the deletion pass
  5. Data: rsyncs posturi.sqlite alone, no --delete and no --inplace, so rsync's write-temp-then-rename swaps it atomically and a request mid-transfer still sees the whole previous database
  6. Code deploy only: stamps the current git rev-parse --short HEAD into webapp-php/static/code-version.txt (gitignored) before the rsync, so the host names the checkout it runs
  7. Checks SITE_URL returns HTTP 200, then verifies the served markers on /versiuni.json — code after a code push, data.built_at after a data push — with up to 5 attempts 5 s apart, since the shared host may serve the old file for a moment. A mismatch fails the deploy. All of step 7 is skipped when SITE_URL is unset

/versiuni.json (webapp-php/feeds/versiuni.json.php, Cache-Control: no-store) reports the deployed code marker and the data provenance (built_at, run_id, git_sha, postings, active_only, detail_fetched_at_max, detail_fetched_rows, index_checked_at) as two separate facts; source_host is not exposed. --verify compares the same two markers against this checkout and the local posturi.sqlite.

Order matters across a schema change. Code and data deploy separately, so new PHP code can land on an export written by older pipeline code. db() covers the known gap: when job_postings lacks application_status (exports before FIX-05), it shadows the table with a TEMP view that computes the status exactly as the export does (shim_legacy_export() in webapp-php/db.php, tested by tests/compat_test.php). Any new column the templates start reading needs the same treatment or a VPS-first rollout: upgrade the VPS checkout, migrate, export and --data-only first, then --code-only.

The full archive stays in PostgreSQL; the deployed SQLite only contains currently active postings.

The export builds beside its target and only replaces it once it opens, passes integrity_check, clears --min-rows (default 100) and has not lost half its rows against the file it would replace. --force overrides the row floors. This is the guard that was missing when the live site served three active postings for five weeks.

Each export carries a build_meta row — when it was built, from which commit and host, and what it holds. FIX-06 added the provenance fields: run_id (the pipeline run, from POSTURI_RUN_ID), index_checked_at / index_scan_pages / index_scan_complete (from data/index-scan.json), and detail_fetched_at_max / detail_fetched_rows (newest Detail Fetched At and how many rows have one); all are empty/null on legacy builds = unknown, never synthesized from build time. The site reads it for the provenance block on /despre/ and for the tooltip behind the "actualizat" stamps; build_meta() in helpers.php returns null on an older export, and every consumer renders without it. Note the two timestamps mean different things: the visible stamp is MAX(last_seen_at), when the source was last scraped, while built_at is when the file was generated. source_host is recorded for the deploy to verify against, never rendered.

job_postings.application_status is the one vocabulary every surface filters on: confirmed_open (deadline from the concurs or the announcement, still ahead), unconfirmed (only the listing's expiry date, still ahead), closed (any date in the past) and unknown (no date). Cancelled postings are excluded from the export rather than given a status.

The export also carries llm_costs — one row per Bucharest day, provider, model and prompt version, summed from jobs_jobpostingschemavariant — which feeds the "Costuri inferență" section of /statistici. It is not cut down by --active-only, and it covers the schema extraction step only: infer and occupations record no tokens or cost. A variant is upserted, so a re-extracted posting's cost moves to its newest day. An export that predates the table makes the section read "Date indisponibile" instead of failing.

Continuous deployment

The pipeline runs unattended on a VPS twice a day (13:15 and 18:33 Europe/Bucharest) and pushes a fresh database to the shared host. The repo ships a systemd timer pinned to Europe/Bucharest (ops/systemd/posturi-pipeline.timer, needs systemd >= 240); the live box still runs a user crontab until OPS-01 is completed. Code deploys stay manual from the development machine, which is why the two rsyncs above are separate — a cron that pushed the whole directory would revert templates from the VPS's older checkout.

Mac (dev) ──git push──> GitHub ──git pull (manual)──> VPS
  │                                                    │
  │  ./deploy-php.sh --code-only     ops/run-pipeline.sh (timer / cron)
  └──────────────> shared host (PHP + posturi.sqlite) <┘
Path What it is
ops/run-pipeline.sh the unattended entry point: flock, pipeline, export check, data deploy, healthcheck ping
ops/check-export.py the deploy gate — asserts the SQLite about to ship is fit to ship
ops/systemd/posturi-pipeline.{service,timer} the two daily slots (not what the live box runs yet; the units hard-code User=posturi and /srv/posturi)
ops/env.sh .env reader shared by the shell scripts (it is never sourced — it holds API keys)
ops/logrotate.posturi rotation for logs/pipeline.log, installed into /etc/logrotate.d/
docs/deploy-vps.md provisioning runbook, operating commands, failure table

Observability

Each run leaves three things behind:

Where What
data/pipeline-runs.jsonl one kind: "run" record (step timings, exit codes, flags) and one kind: "export-check" record (33 data metrics, every assertion), joined by run_id; plus kind: "llm-balance" (the pre-flight balance and its outcome)
logs/pipeline.log the steps' own output, runs delimited by === posturi pipeline run <id> … ===
HEALTHCHECK_URL /start, /fail, success — the only signal that can report a run which never happened

By default a failed pipeline step blocks the deploy: ops/run-pipeline.sh pings /fail, exits with the pipeline's status and the host keeps serving the previous database. Set POSTURI_ALLOW_DEGRADED_DEPLOY=1 to deploy anyway when the export check passes (the run is recorded as degraded).

ops/check-llm-balance.py runs first, before anything is fetched or paid for: it reads the DeepSeek balance (GET /user/balance, free) and exits 69 below LLM_BALANCE_MIN (default 0.50 USD, or when the account reports is_available: false), which ops/run-pipeline.sh turns into a /fail ping and an immediate stop. Below LLM_BALANCE_WARN (default 3.00) it only warns. A check that cannot reach the API, or a provider other than DeepSeek, prints a line and exits 0 — it never blocks a run. POSTURI_SKIP_BALANCE_CHECK=1 bypasses it; --need USD is for a manual backfill of known cost. Exit codes of ops/run-pipeline.sh: 0 fine, 65 export failed its hard checks, 69 balance too low, 75 another run holds the lock, 78 unapplied migrations.

ops/check-export.py runs between the export and the deploy. Hard checks abort before the rsync, so a corrupt export cannot reach the live site: integrity_check, broken foreign keys, duplicate URLs, a short FTS index, anything other than exactly 42 județ rows, a collapse below half the previous run, or a build_meta.built_at older than six hours (which means the export step failed and the deploy is about to ship an old file). Soft checks — coverage regressions, an altele spike, one profession family or skill tag taking an implausible share, anomaly-flag step changes, mojibake — are recorded and printed but do not block; --strict makes them fatal.

Every check is tied to a regression that actually happened here, and the two *_sanity_warnings() functions from import_csvs.py are called rather than restated, so "too many counties" has one definition in the tree.

.venv/bin/python ops/check-export.py --no-log              # read the live numbers
.venv/bin/python ops/check-export.py --db /tmp/copy.sqlite --no-log

--no-log keeps an interactive read from becoming the baseline the next real run compares against. For the narrative pass over the run log — missed slots, recurring step failures, metrics sliding while every individual check still passes — run /pipeline-check in Claude Code. See docs/pipeline-quality-checks.md for the full watch-list and what is still unbuilt.

Full setup — packages, seeding Postgres and the scrape cache from the Mac, SSH keys, installing the timer — is in docs/deploy-vps.md.

Stylesheet

The app has no build step on the host — the compiled CSS is committed. Rebuild it whenever you touch a template's classes:

npm install          # once
npm run css          # webapp-php/assets/app.css -> webapp-php/static/app.css (minified)
npm run css:watch    # while editing templates

Fonts (static/fonts/*.woff2) and htmx (static/htmx.min.js) are self-hosted; nothing is fetched from a CDN at runtime.

Skins

The whole palette is one :root block of CSS custom properties in webapp-php/assets/app.css. Tailwind's theme resolves every colour, radius and font utility to one of those variables, so re-declaring them under [data-skin="<id>"] restyles the entire site without touching a single utility class. That is all a skin is.

Three ship today, chosen with the picker in the footer (stored in localStorage, applied before first paint so there is no flash):

id what it is
hartie the built-in default — parchment, Fraunces, DM Sans. No file; it is :root.
govuk GOV.UK Design System — white, square, Arial, yellow focus, black masthead.
posturi the official posturi.gov.ro — navy + gold, Manrope, rounded, lifted cards.

Adding one: copy webapp-php/static/skins/_template.css to <name>.css and it appears in the picker on the next request — no registry, no build step. The id is the filename; inc/skins.php discovers it and reads the display name from the @skin comment. Files starting with _ are skipped.

Two rules the template explains at length:

  1. Scope every rule under [data-skin="<id>"]. All skin files load on every page, so an unscoped rule leaks into the others. @font-face is the exception — it declares a family rather than applying one, and is how a skin ships its own typeface.
  2. Colours are space-separated RGB channels (245 240 232), not hex. That is the only form Tailwind's alpha modifiers compose with; a hex value silently breaks every bg-surface/70 on the page.

Skins live in static/ rather than assets/ because they need no build — Tailwind would strip them, since nothing in the PHP references their selectors — and because deploy-php.sh excludes assets/.

Validate before committing:

php webapp-php/assets/check-skins.php

It reads the token contract out of app.css and flags the four failures that are silent in a browser: an unscoped rule, a token name that does not exist, a colour written as hex, and any text/background pair under WCAG AA.

Previewing another export

db.php honours a POSTURI_DB environment variable, so a test or a preview can point at a different SQLite file without touching the deployed one:

POSTURI_DB=/tmp/other.sqlite php -S localhost:8000 -t webapp-php webapp-php/router.php

Unset in production.

Local preview

php -S localhost:8000 -t webapp-php webapp-php/router.php

router.php exists only for PHP's built-in server — it serves static files and mirrors the .htaccess deny rules. Apache never loads it.

Structure

Path Purpose
index.php Front controller — routes /, /job/123-slug/, /angajatori/, /statistici/, /despre/, /robots.txt, /sitemap.xml
db.php PDO singleton for posturi.sqlite
helpers.php Markdown rendering, date formatting, filter builder, facet queries, FTS query builder, active-filter chips, URL helpers
pages/ List, detail, employer list, employer detail, stats, about
feeds/ Atom, JSON API, iCal endpoints — follow the active filters; ?employer=<slug> scopes a feed to one angajator (linked from each employer profile)
partials/ Result list and facet sidebar (both HTMX out-of-band swappable)
inc/ Header/footer HTML, <head> metadata (canonical, OpenGraph, JSON-LD hook)
assets/ Tailwind source — not deployed
static/ Compiled app.css, self-hosted fonts, htmx
router.php Dev-server router — not deployed
.htaccess URL rewriting, asset caching, blocks direct access to *.sqlite*

About

posturi.gov.ro, da mai drăgu'

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Contributors

Languages