Engineering · transparent
Methodology
How OpenCaseLaw finds, weights and verifies Swiss court decisions — end to end, with every weight and threshold disclosed. This page is the authoritative reference; what you read elsewhere should agree with what is recorded here — otherwise the code wins.
1. Corpus & updating
The corpus comprises 1,050,000+ published Swiss decisions from 118 courts, in three languages (DE 499,642; FR 459,825; IT 90,776), spanning 1875 – 2026. Eight federal courts (BGer 192,049; BVGer 108,135; BGE 50,459; plus BStGer, BPatGer, MKG and others), 26 cantonal jurisdictions, and ECHR-Switzerland (~2,800 decisions from BGE translations, HUDOC, Chamber/Grand Chamber/Committee).
Each decision is sourced from its primary source — where available, from the court's own publication portal (all 26 cantons fetched directly via LexWork / SIL / ZH OpenData / TI RL, with a LexFind PDF fallback that fills missing enactments in 4 cantons), and for federal updates from the BGer's official search back-end. 40 of 51 entscheidsuche.ch shards are mirrored by our own direct scrapers; entscheidsuche data only prevails in deduplication where we have no live equivalent of our own. The remaining seven entscheidsuche-exclusive shards (≈142,000 records: vd_findinfo, vd_omni, ch_vb, sg_gerichte, be_bvd, be_weitere, be_steuerrekurs) are frozen historical archives — their upstream portals have been retired, replaced or temporarily disconnected. We re-ingest entscheidsuche weekly as a completeness safeguard; the last six weekly runs produced zero new decisions. See Coverage → supplementary archives for the full breakdown. The pipeline updates on three schedules:
- Every 15 minutes on weekdays from 05:00–16:00 UTC — the BGer poller publishes new federal decisions incrementally into the live database, within minutes of their court-side release.
- Daily at 01:00 UTC — every active cantonal scraper runs; a soft-fail policy is tied to the 5 critical federal scrapers and a 15 % error-rate threshold across the rest.
- Daily at 03:30 UTC — full FTS5 rebuild, citation-graph rebuild, quality gate, statistics update, zero-downtime atomic swap. The workers keep serving the existing database until the new one has passed its integrity check.
The full corpus is also published as Parquet on Hugging Face (voilaj/swiss-caselaw) under CC0; the daily delta is appended automatically.
2. Full-text index
SQLite FTS5 with the unicode61 remove_diacritics 2 tokenizer. The decision table feeds the index through these columns, with per-column BM25 weights tuned empirically against a 100-query golden set:
| Column | BM25 weight | Rationale |
|---|---|---|
| title | 6.0 | Highest signal — terse, deliberate. |
| regeste | 5.5 | The court's own thematic summary. |
| docket_number | 2.0 | For exact retrieval by docket number. |
| full_text | 1.2 | Long, noisy — anchor, not driver. |
| court / canton / language / decision_id | 0.8 | Metadata, not content. |
The index is rebuilt atomically: build_fts5.py writes to decisions.db.tmp and then replaces the file via os.replace(). Workers using the SQLite URI ?immutable=1 keep their open file handles on the old inode until they reconnect — the swap stays invisible to active readers.
build_fts5.py · BM25 weights configured in mcp_server.py around lines 431–442 · diacritic tokenizer mirrored in decision_structure.db for paragraph-level search.
3. Query understanding
A natural-language query never reaches FTS5 unprocessed. It first passes through sanitisation (_sanitize_fts5): apostrophes, hyphens and dots without a word character are reduced to spaces; reserved tokens (AND, OR, NOT, NEAR) are kept only when they have operands on both sides; the single Swiss legal token "OR" (abbreviation for the Code of Obligations) is always forced into quotation marks, as it would otherwise be interpreted as a Boolean operator.
In parallel, the query is routed to Claude Haiku 4.5 for a structured 2-second analysis: statutory references, doctrinal terms in DE/FR/IT, mentions of authoritative BGE, synonyms and the legal area are extracted as deterministic JSON. The output is cached by the lowercased query for the duration of the session.
Three normalisations make the Swiss legal vocabulary searchable across its stylistic, orthographic and cross-lingual variants:
- Diacritic-insensitive tokenisation. The FTS5 tokenizer
unicode61 remove_diacritics 2strips diacritics on both the index and the query side — so "Prüfung", "PRÜFUNG" and a query for "Prufung" all hit the same result list. - Umlaut-spelling merging. The diacritic-free spelling
ae/oe/ue(common in older judgments and in any pre-Unicode environment) is reduced toa/o/u, which the tokenizer then unifies with the diacritic-strippedä/ö/ü. So "Pruefung" also finds "Prüfung". - LLM synonym expansion. Claude Haiku's structured analysis produces 2–4 alternative legal terms in DE/FR/IT per query, on the fly — not from a static table. So "qualité pour recourir" is linked to "Beschwerdebefugnis" / "Beschwerdelegitimation" / "legittimazione", even if your query contained only one of them.
4. Retrieval & fusion
The query fans out into 10–12 strategies that run as independent FTS5/graph queries. Each strategy carries a strategy weight; one strategy's rank-1 hit and another's rank-1 hit are fused via Reciprocal Rank Fusion (RRF rank constant = 60).
| Strategy | Weight | Effect |
|---|---|---|
| nl_and | 1.8 | Natural-language AND over the terms. |
| raw | 1.5 | Query verbatim. |
| regeste_focus | 1.4 | Restricted to the regeste column. |
| nl_or | 1.2 | OR fallback (cost-aware early termination). |
| structured_doctrine | 1.1–3.5 | Doctrinal terms from the Haiku analysis. |
| quoted_explicit | 1.1 | Phrase match when quotation marks are detected. |
| nl_or_expanded | 1.0 | OR + synonym / umlaut / compound expansion. |
| title_focus | 0.95 | Match restricted to the title column. |
| doctrine_regeste / doctrine_title | 2.5 / 1.6 | Term-translation strategies. |
Beyond FTS5, the same RRF pool also receives:
- Candidates from the statute graph — decisions linked to the statutory articles parsed from the query, fed in as a parallel ranking.
- Direct BGE look-ups — if the query (or its analysis) names a specific BGE reference, that decision is pinned to the top.
- Docket-number matches — strings like
6B_1234/2025are routed to an exact docket-number look-up, which skips most of the pipeline.
The fused candidate pool is sized dynamically (by default ~300–400 for a top-50 query) and capped at 2,500 rows before re-ranking begins.
5. Re-ranking
Each candidate receives a vector of signals; the final score is a linear combination tuned against the golden set. Signal weights (_rerank_rows):
| Signal | Weight | Caps / notes |
|---|---|---|
| RRF score | 32.0 | Aggregate of all strategy ranks. |
| Docket number exact / partial | 6.0 / 2.0 | Match at the string level of the docket number. |
| Title coverage | 3.0 | Share of the query tokens in the title. |
| Regeste coverage | 3.0 | The same, against the regeste. |
| Statute mentions | 3.5 · 0.5 · 2.0 | Decision-to-article links. |
| Citation matches | 2.4 · 0.30 · 1.2 | Cross-result citation evidence. |
| Authority (incoming citations) | 0.03 · 1.0 | Why a leading case rises. |
| Language match | +2.0 | When the result's language matches the query. |
| Expanded coverage | 1.5 / 0.8 | Credit for synonym + compound matches. |
| Court-domain heuristics | ±0.2 – +1.7 | BVGer-asylum + BGer-supreme-court weighting on matching intent. |
After the linear scorer, an LLM re-ranking pass is triggered when needed:
- Model: Claude Haiku 4.5; Top-N: 15; timeout: 3 s; weight: 3.0 with linear decay
w × max(0, 1 − rank/15). - A confidence gate skips the call entirely when the lexical top-1 score is already ≥ 2× the top-2 score (the answer is unambiguous; re-ranking is pure cost).
- The pass is also skipped for docket-number queries (an exact match needs no LLM).
6. Pinpoint references live · May 2026
Each of the top five search results and each of the top three leading-case results carries a pinpoint field that names the consideration most relevant to the query (paragraph holding the legal decision). The resolver runs over a per-decision FTS5 paragraph index in decision_structure.db (≈8.8 M paragraphs across 807,000 decisions). The two-stage design:
- Phrase pass. The query is run as an exact FTS5 phrase — high precision when the user's wording matches the court's.
- Bag-of-words OR pass. If the phrase returns nothing, the same tokens are fired as an OR query, for broader retrieval.
A confidence scorer (_score_pinpoint_confidence) combines three independent signals to flag the result:
- BM25 gap. Rank-1 score over rank-2; a ratio > 1.5 yields "high", > 1.2 yields "medium".
- Absolute strength. For single-row matches with no rank-2 to compare against: absolute |BM25| > 2.0 for "high", > 1.0 for "medium". An earlier version of this branch used a sentinel value of 999.0 for single-row matches, which silently promoted thin matches to "high"; that was the false-confidence bug fixed in May 2026.
- Token coverage. Unique query tokens (> 2 characters, with ~70 generic Swiss legal-discourse stopwords filtered out) appearing in the matched paragraph. Multi-token queries with < 50 % coverage are suppressed entirely; < 70 % is capped at "medium".
Semantic rescue infrastructure rolled out · corpus ~33 % encoded kicks in when the lexical pass returns nothing. The query is encoded with sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 (117 M parameters, 384-dimensional normalised vectors, multilingual DE/FR/IT/RM plus 47 more); cosine similarity is computed over the decision's < 300 paragraph embeddings; a result appears as "high" if cosine ≥ 0.70, "medium" at ≥ 0.55, otherwise none. A hybrid mode, enabled once encoding completes, runs both signals and treats agreement on the same consideration as cross-signal evidence — the confidence is raised to "high", source: "hybrid_agreement".
The pinpoint URL is built so that the browser automatically scrolls to the matched paragraph and highlights it: ?highlight=<verbatim sentence>&e=2.3#e-2-3. The query parameter drives a server-side highlight; the hash fragment triggers the browser's native auto-scroll.
7. Citation graph
Each decision's text is parsed for citations of other decisions (BGE references, BGer docket numbers, BVGer references). The raw layer holds 9.65 M edges; the resolver lifts 8.09 M of them (92.9 %) to canonical decision IDs. The remaining 7 % are real decisions we do not yet hold (older pre-1875 BGE, withdrawn court publications, OCR-corrupted references).
Resolution runs as four cascaded passes:
- Standard docket-number join. Normalised docket number of the citation = normalised docket number of the target. Handles over 80 % of the edges.
- BGE prefix bidirectional. Finds matches regardless of whether the citation contains the
BGEprefix, with the target stored in either form. This single March 2026 fix nearly tripled BGE resolution. - Bare BGE numbers. Targets matching
volume division pageare treated as bare BGE references and prefixed withBGE. - Pin-cite resolution. If a citation contains a page number absent from our targets (e.g.
BGE 125 V 352, where the first page is 351), we look for the largest first page ≤ 352 within the same (volume, division) and within 30 pages. Confidence is lowered by 0.10 to reflect the inference.
Most-cited authority (computed live from reference_graph.db on 2026-05-11): BGE 125 V 351 with 85,108 incoming citations, followed by BGE 134 V 231 and BGE 122 V 157, each in the tens of thousands. These counts feed back as an authority signal in search re-ranking (weight 0.03 per citation, capped at 1.0) — which is why classic leading cases rise, even when their linguistic matches are no stronger than other candidates'.
8. Quality assurance
Every nightly publication is guarded by a 4-layer QA framework. L4 (the publication gate, which runs the CRITICAL subset of L1) is the hard lock: if it fails, the swap is refused and users keep yesterday's corpus until the issue is investigated. The other layers are continuous safeguards — L2 runs on every commit in CI, L3 every 5 minutes on the live server.
- L1 — Dataset checks (63 across 20 modules). Per-court drift, duplicate detection, short-text / OCR detection, citation-graph resolution rate, statute-link coverage, date / docket-number plausibility, missing-field counts, plus LLM spot-checking on a rotating selection.
- L2 — pytest suite (currently 516 passing). Unit tests cover every parsing, dedup and ranking primitive; new regressions land as new tests following the "incident → regression test" pattern.
- L3 — Smoke test (every 5 minutes). Production health probe:
/health, an anchor decision page and the PDF export pipeline. Three consecutive failures escalate to an INVESTIGATE alert. - L4 — Publication gate. Of the L1 checks, those marked CRITICAL must pass before the new corpus is committed and pushed; if one fails, publication refuses the swap and yesterday's data stays live. The dashboard at /quality.html shows the latest status and history of every check.
Beyond QA, decision-citing LLM workflows pass through a five-track final audit (attest_response) that catches four hallucination classes — fabricated citation, fabricated quotation, fabricated statute text, fabricated date — plus an optional fifth grounding judge, which checks whether the claim the LLM attributed to a cited paragraph is actually supported by that paragraph's text.
9. Tested and rejected
Technical decisions become more trustworthy when the dead ends are visible. The following techniques were measured on the same golden set as the rest and did not earn their place:
- Cross-encoder re-ranking (
cross-encoder/mmarco-mMiniLMv2-L12-H384-v1as the current placeholder) rejected Multilingual cross-encoders trained on generic web QA did not transfer cleanly to the Swiss legal vocabulary on the golden set — they hurt MRR instead of helping. The helper remains in the code behind_apply_cross_encoder_boostsfor future fine-tuning experiments, but is not called in production. - Pre-built BGE-M3 dense vectors rejected The corpus encoded, vector RRF added as a fourth strategy. No improvement over BM25 + RRF + Haiku at any vector weight. The
vectors.dbwas removed in March 2026. The per-paragraph semantic embeddings shipped this week are a separate, narrower intervention — they only score within a single decision's paragraphs, where BM25 has well-known limits. - Larger Haiku top-N rejected Re-ranking the top-30, top-50 instead of the top-15 brought no MRR gain and a linear cost increase. Top-15 is the empirical sweet spot for confidence-driven re-ranking.
- Larger candidate pool rejected Enlarging the pool beyond ~400 candidates (for standard queries) did not move recall meaningfully; the FTS5 + RRF combination already surfaces the right candidates within the first few hundred results. The 2,500 limit remains as a safety ceiling.
10. Open access
The corpus is CC0 (public domain dedication). The code is MIT, hosted at github.com/jonashertner/opencaselaw — every weight, threshold and heuristic on this page is auditable in source. No accounts, no cookies; what the server records, including search queries, and for how long, is documented in the privacy policy: /datenschutz/.
Programmatic access:
- REST API. mcp.opencaselaw.ch/docs — OpenAPI 3.0.3 with 47 endpoints (search, get, leading cases, trends, citation graph, statutes, commentaries, Botschaft, exports, pinpoint, attestation).
- MCP server. mcp.opencaselaw.ch/sse — Server-Sent Events transport, exposing 42 tools. Connected by Claude, ChatGPT (OpenAI MCP), Cursor, Microsoft Copilot Studio, Perplexity, and others.
- Hugging Face dataset. voilaj/swiss-caselaw — Parquet shards, daily delta-published.
- Word add-in. opencaselaw.ch/word/ — Office add-in for Microsoft Word; surfaces the same search + pinpoint stack inside the document.
This page reflects the live system. If you spot a discrepancy between what's described here and what the source says, the source wins — open an issue and we'll fix one of them.