Privacy

Privacy

OpenCaseLaw logs requests to operate the service and to support research, but sets no cookies and runs no third-party tracking. Every published number is aggregated, k-anonymous, and differentially private.

§01 Summary

The short version

OpenCaseLaw is a public research project. There is no user account and no persistent identifier that automatically links a request to a person. That is not a claim that we collect no data: query text, timestamps, session and client identifiers, and server access logs are all recorded, as set out in full below. This page states what is actually collected, why, and for how long.

  • No cookies. The site sets no cookies at all, neither necessary nor optional.
  • No fingerprinting. No canvas, font, audio, or WebGL fingerprinting. No third-party analytics.
  • No Google Analytics, no Plausible Cloud, no Sentry. All analysis runs on our own server in Germany.
  • Development data: full capture, permanent research archive. To improve search, for research and to train our own models, we record complete request data: tool, parameters including query text, timestamp, client class and a session identifier. No IP address is stored with any request; sessions cannot be linked to persons — there is no account, no login, and the identifier rotates with every connection. Document text from the Word add-in is never stored, only its length. Query texts are retained indefinitely, stripped of session and personal references, in an access-restricted private research archive — for search quality, research and model training; no time limit applies to it. How long raw records carrying a session reference stay on the serving systems is at the developer’s discretion. What persists permanently is aggregate statistics, this query archive, and models trained from them. The collection code is open source (mcp_server.py, scripts/collect_dev_data.py).
  • No user account, no persistent identifier. There is no login. Identifiers that do exist are time-limited and rotate: the MCP session id with every connection, the cost-ledger pseudonym daily, the Word add-in install-cohort hash monthly (details below, §02 and §04).
  • All published numbers are differentially private (ε = 1.0) and subject to a k-anonymity threshold of k = 10.

§02 Logs & retention

Server logs in three tiers

Every web server logs requests — that is technically unavoidable. Part of our logging decays into harmless aggregates within days; the exceptions with longer or indefinite retention are below.

Tier 1 — access.log (72 hours)

Contains IP address and User-Agent. Legal basis: legitimate interest in defending against DDoS and abuse. Retention: strictly 72 hours, then cryptographically shredded (shred -u). Never used for analytics.

Cost and abuse control

Two separate systems are at work here. First, per-IP daily limits for LLM-backed endpoints (attest, verify-claim, check_claim_support and others): IP address, endpoint and call count per day are recorded. No automatic deletion currently applies to these counters. Second, a per-call cost ledger: it holds a daily-rotating pseudonym of the IP address (never the address itself), the tool called, and the cost of the call. Retention: 90 days on the serving systems; in addition we keep an access-restricted private archive of these entries for the lifetime of the service. Purpose of both systems: daily limits, cost control, defence against and documentation of commercial abuse. Higher limits: team@jonashertner.com.

Search-quality traces (30 days)

To measure ranking quality (BM25, cross-encoder, Haiku rerank) we log candidate ids and timings per search; Haiku-rerank entries additionally carry the first 200 characters of the query text. This record holds no IP address and no session id; the result_set_id it carries can, when full capture (above) is on, be joined to a session-linked fetch event from the same day. Retention: 30 days, then deleted automatically.

AI processing via the Anthropic API (per request)

Several functions send text to Anthropic's API (api.anthropic.com, Claude models) for processing at request time: query parsing, query expansion and re-ranking of search results send the query text and short excerpts of the candidate decisions; check_claim_support sends the submitted claim together with the passage it is checked against; attest_response with audit_grounding=true sends the sentences around each citation together with the cited Erwägung; the Word add-in's Verify, Ground and Mirror features (/billing/verify, /billing/find-support, /billing/reflect) send the selected text or the document text; Audit (/attest) and Strengthen stay on our server. All five Pro endpoints receive the text only after the add-in has replaced nine categories of personal data with placeholders (e-mail, AHV number, IBAN, UID, phone, date of birth, address, postal code/city, titled names; pattern-based, no guarantee of completeness — scope and limits at word.opencaselaw.ch/privacy.html); the server re-checks the five structured patterns and otherwise rejects the request. Never sent: IP address, session id, client class, or any identifier. The transmission is subject to Anthropic's commercial API terms and privacy policy; retention at Anthropic is not under our control. On our side we record per call the model, token counts and cost (see cost ledger above), and for re-ranking the first 200 characters of the query (see search-quality traces). Legal basis: performance of the requested function; these functions cannot run without this processing.

Tier 2 — tier2.log (14 days, no personal data)

Contains only class labels: client class (e.g. “cursor”, “claude_hosted”, “word_addin”), endpoint class (e.g. “rest_search_decisions”), HTTP status, response time, response size. No IPs. No User-Agents. No query strings. No referer. This file is structurally incapable of identifying a person. Retention: 14 days as a safety net, then deleted.

Tier 3 — analytics.db (forever, aggregated & private)

Daily aggregates per (client class, endpoint class). No per-user rows, no finer-than-daily granularity. Cells with n < 10 are suppressed in the public column. Published numbers carry Laplace noise at ε = 1.0 (formal (ε, 0)-differential privacy). This file may be and will be published publicly — because it is mathematically proven to reveal nothing about individuals.

Unique-user counts (counting sketch)

To determine daily and monthly numbers of unique direct consumers, IP addresses are hashed in memory only, with a rotating secret, and fed into a HyperLogLog counting sketch. Only sketch registers and totals reach disk; they mathematically cannot reveal whether any particular address was present. The secret is deleted at each day and month boundary, which also rules out any after-the-fact linkage across windows.

Microsoft 365 Copilot agent (/mcp-edu)

The OpenCaseLaw agent for Microsoft 365 Copilot calls its own endpoint, /mcp-edu. For requests through this endpoint there is no full capture (no query text, no session id, no research archive, no model training), no search-quality traces and no per-IP cost ledger. What remains is the tier-1 log (here with Microsoft's IP address, 72 hours), tiers 2 and 3 without personal data, and cost accounting without IP address or query text (model, tokens, cost). Query analysis and re-ranking via Anthropic and the search of cantonal legislation via LexFind work as described above.

The rollup code is open source and auditable: scripts/rollup_analytics.py. The nginx three-tier logging config is at ops/nginx/ocl-logging.conf.

§03 Aggregation & DP

What we measure, and how

We measure only what is necessary for operational decisions — and we measure it at aggregate level, never per user. The following metrics are computed nightly and published at opencaselaw.ch/stats.json:

  • Tool calls per day, broken down by client class and endpoint.
  • p50 and p95 response time per endpoint.
  • Error rate per endpoint (4xx, 5xx, 429).
  • Approximate active installations per client class (via a HyperLogLog sketch that rotates monthly).
  • HTTP status distribution across all requests.

What we never collect

  • Permanent individual profiles — no IP address with requests, no user account, no link between a session and a person. During the evaluation period, requests within the same session are linked to one another; how long those raw records are kept is at the developer’s discretion and follows the stated purposes. What is kept permanently carries no session or personal reference.
  • Which decisions or laws a specific person looks up.
  • Raw IP addresses in aggregates or Tier-3 data.
  • Full User-Agent strings beyond 72 hours (only the class).
  • Referer header — would reveal which chat or document a tool was called from.
  • Fingerprinting (canvas, fonts, WebGL, audio) — never, under any circumstances.
  • Linking Stripe customer data (Word add-in Pro) to usage data.

§04 Word add-in

Word add-in — a special case

The Word add-in is the only part of OpenCaseLaw that optionally sends an install cohort hash so we can estimate the approximate number of active installations per month. The mechanism is deliberately built so that cross-month tracking is cryptographically impossible:

  1. On first run, the add-in generates a random UUID and stores it locally in localStorage.
  2. On every API request it computes SHA-256(uuid + "YYYY-MM") and sends the first 8 hex chars as an X-Install-Cohort header.
  3. The monthly salt changes each month. April and May hashes are cryptographically unlinkable — even if we kept all hashes forever (which we don’t).
  4. We use these hashes only to fill a daily HyperLogLog sketch per client class. The sketch yields an estimated count — the input values are not recoverable from it.

Opt-out: in the add-in menu under “Settings → Anonymous usage signal” you can disable this header at any time. The add-in works identically without it.

Pro subscriptions are handled by Stripe. OpenCaseLaw sees only the license key and matches it during Pro requests. There is no link between Stripe customer data and usage statistics.

§05 Your rights

Your rights (nFADP / GDPR)

The revised Swiss Federal Act on Data Protection (nFADP, in force since September 2023) and the European GDPR grant you rights of access, rectification, and deletion. OpenCaseLaw keeps no user account and no persistent identifier that automatically links a request to a person. That limits what we can attribute to you in practice, but it does not remove these rights:

  • Right of access: without an account or persistent identifier, we generally cannot connect a request to you on our own. If you give us your session id, an approximate time and the tool used (for example from your own client logs), we search the systems described in §02 for matching entries and tell you what we find.
  • Right to deletion: entries that can be attributed to you this way are deleted on request, to the extent deletion is technically possible in the system holding them. If you installed the Word add-in, uninstalling it additionally removes the local install UUID from your device.
  • Portability: where entries can be attributed to you, we provide them on request in a common format.
  • Objection: you can object to the 72-hour Tier-1 logs by not using the services; an exception is technically impossible, because web servers must log requests in order to respond to abuse. For the other systems described in §02, send your objection to the contact address below.

Data controller under the nFADP: Jonas Hertner, email team@jonashertner.com. Servers are operated by Hetzner Online GmbH in Germany.

§06 Contact

Questions, concerns, audits

Every mechanism described here is open source and auditable in the GitHub repository. If you spot a discrepancy between this page and the system’s actual behaviour, treat it as a bug — report it and we will fix it.

Email: team@jonashertner.com
Privacy questions get a reply within a few days.

Last updated: 2026-09-25 · Versioned in Git, changes publicly visible.