Full-text (BM25) search for codified law (statutes, regulations, constitutional provisions) matching keywords within a jurisdiction.
This is the unified search tool for the v0 Surveyor surface. Scope is controlled by the optional `law_key` and `within_division_path` params:
- **No `law_key`**: searches across ALL law types in the jurisdiction (statutes, regulations, constitutions). Returns a single ranked list. Best for natural-language questions when you don't know which law type holds the answer.
- **`law_key` set**: scopes to one specific law (e.g. `CA-STAT`, `TX-RR`, `FED-CFR`). The RELIABLE way to sidestep BM25 ranking biases that can cause one law type to crowd out another in jurisdiction-wide search. Required when you want to drill into a `within_division_path`.
- **`within_division_path` set**: scopes the search to a subtree of the law (e.g. `title_15.division_1.chapter_1`). Requires `law_key`. Useful for iterative drill-down. Pass a list to search across multiple sibling subtrees in one call (OR semantics, comma-joined under the hood).
- **`with_federal=true`**: also include federal laws in the search alongside the named state jurisdiction.
- **`query_type` set**: switches the matching mode. Defaults to the API's default (whitespace-split, every word must appear). Set to `"phrase"` as a tokenization-gap workaround when the default mode returns thin results on apostrophe- or punctuation-containing queries — see TOKENIZATION FALLBACK below.
Results are returned as LEAN SUMMARIES — `path`, identifier, display_name, display_ancestors (context; may be empty), lifecycle flags, and dates. `display_ancestors` is a list of display-name strings (root-most first), e.g. `["Title 28: Transportation", "Chapter 23: Highway Beautification", "Article 1"]` — NOT a list of dicts. To get the full text of a winning result, pass its `path` to `get_division_by_path` (with the result's `jurisdiction_key` and `law_key`) — `path` is always present and needs no citation formatting, so this is the reliable retrieve step. Alternatively, if you have a clean citation, build a strict Bluebook form from `identifier` (plus `display_ancestors` when present), e.g. `Tex. Lab. Code § 406.033`, and call `resolve_citation`. Either way the workflow is: search, skim results, retrieve primary source, quote.
COMMON law_keys (discover the full list via `list_jurisdictions`):
- State statutes: `CA-STAT`, `NY-STAT`, `TX-STAT`, `FL-STAT`, ...
- State regulations: `CA-RR`, `NY-RR`, `TX-RR`, `FL-RR`, ...
- State constitutions: `CA-CONST`, `NY-CONST`, ...
- Federal statutes: `FED-USC`, `US-USC`
- Federal regulations: `FED-CFR`, `US-CFR`
IMPORTANT — BM25 scoring can favor regulations over statutes for natural-language queries. Regulations have longer, more keyword-dense operational text, so they often outrank shorter statutory provisions even when the statute is the better answer. If the results look homogeneous (e.g., all from one law_type like CA-RR) but you suspect the right answer is a statute (or vice versa), do one of:
1. Re-call this tool with `law_key` set to the expected law_type (e.g. `CA-STAT` for California statutes) to scope the search and bypass the ranking bias. Reliable move.
2. Rephrase the query with vocabulary closer to the target law type's lexicon. Statutes tend to use short, declarative language ("shall notify", "is guilty of", "harbors conceals aids", "felony") rather than operational language ("duty to report", "mandatory reporting requirements"). If your first query reads like a regulation's prose, try rewriting it to sound more like a statute.
3. If neither surfaces the expected answer, honestly consider that the relevant provision may not be in the corpus OR the query vocabulary may be too far from the text. "No responsive statute found" is a valid answer; confidently returning adjacent regulations as if they answered a statute question is the failure mode to avoid.
TOKENIZATION FALLBACK: by default this search splits the query on whitespace and requires every word to appear (the API's default mode, which is the "and" tokenizer). If your query contains apostrophes, possessives, or punctuation (`workers' compensation`, `employer's notice`) and the default search returns thin or zero results, that's often a tokenization mismatch, NOT a corpus gap. Some OpenLaws law titles use curly apostrophes (`workers'`) that don't match the straight-apostrophe form (`workers`) under the default tokenizer. Before concluding the topic isn't covered, retry with `query_type="phrase"` — the phrase tokenizer handles apostrophe/punctuation boundaries differently and frequently surfaces results the default mode misses.
Three `query_type` values are supported:
- `"and"` (same as default, None): every word must appear.
- `"or"`: any word may match. Useful for synonym-spread queries; behavior may overlap with `"and"` depending on Vespa's scoring on the specific query.
- `"phrase"`: words are matched as an exact phrase AND the phrase tokenizer is more permissive on apostrophes/punctuation. The primary use is the tokenization workaround above.
HONEST-GAP BEHAVIOR: if this tool returns zero results AND you have strong reason to believe the corpus should contain something relevant, rephrase the query (see BM25 guidance), try `query_type="phrase"` (see TOKENIZATION FALLBACK), or scope to a more likely law type. If broad, scoped, AND phrase searches all come back empty, honestly report "no responsive provision found in the indexed corpus" rather than falling back on tangentially related results.
Defaults to the top 10 results; can request up to 50.
search_codified_law