Invarra field notes
Clear thinking about AI behavior under variation.
Research-led essays on measurement validity, semantic evaluation, deployment evidence, and why a system that is right once may still be unreliable.
All notes
Research, made readable.
The Row Is Often The Wrong Unit
Semantic evaluation often starts with a row. Canonical semantic units give evaluators a better unit of analysis when the same meaning has several surface forms.
Disagreement Is Data
When valid representations of the same underlying case produce different outcomes, disagreement may not be noise. It may be the measurement signal.
Valid Variation Needs A Contract
Semantic-preservation contracts make variation interpretable by stating what must stay fixed, what may change, and how validity is checked.
The Right Null Hypothesis For Indirect Observation
When a target cannot be observed directly, evaluators should assume observed behavior may be representation-sensitive until valid variation supports a stronger claim.
Semantic Brittleness Should Be Attributable
A useful evaluation should not only say that a semantic system is brittle. It should help locate where the brittleness enters the measurement stack.