AI Safety is Reinventing the Law

AI safety has spent a decade re-deriving the conceptual inventory of law.

Legend

Crosswalk

AI safety termLegal termRelationNote
alignmentagency lawsamePrincipal-agent problem under incomplete specification.
operatorprincipalsame
intent alignmentduty of loyaltysameThe agent tries to serve the principal’s ends; whether it succeeds is separate.
capabilitiesduty of caresameFiduciary law splits loyalty from competence; Christiano’s definition draws the same line independently.
value alignmentnatural lawanalogyNorms sourced above any particular principal or sovereign.
outer alignmentlegislative draftingsameWriting text so the text picks out the intent.
inner alignmentsubdelegationanalogyDelegatus non potest delegare. The trained policy is a sub-agent that may not carry the mandate.
mesa-optimizersub-agentanalogy
prompt injectionforged instructionsameBusiness email compromise with the employee swapped for a model.
reward misspecificationincomplete contractsameHadfield-Menell and Hadfield (2019) state the identity outright.
specification gamingsubstance-over-form doctrinesameLetter satisfied, purpose defeated; the doctrine looks through the form.
reward hackingdefeat devicesameVolkswagen detected the test and optimized for it.
Goodhart’s lawrule gamingsameTax avoidance is the canonical instance: optimize the measure, lose the target.
jailbreakloopholesame
alignment taxcompliance costsame
constitution (Constitutional AI)constitutionportSuperordinate norms constraining downstream rule-making. Anthropic chose the word for the function.
model specstatuteportEnacted text, general application, enforced by training rather than courts.
system promptstanding ordersport
content policyterms of servicesameIt is a contract.
chain of commandhierarchy of normsportSupremacy ordering: platform over developer over user. Kelsen without attribution.
RLHFcommon lawanalogyPolicy accreted from judged instances rather than enacted text.
LLM-as-judgejudgeport
red-teamingadversarial processsameTruth-finding by paid opposition.
honeypot evalsting operationsameInducement invalidates the result in both.
capability evallicensure examinationportCapability tested before practice is permitted.
model cardmandated disclosureportA prospectus for a model.
chain-of-thought monitoringrecord requirementanalogyReasoning on the record so review is possible.
interpretabilityreasoned-decision requirementanalogyAdministrative law voids decisions that cannot give reasons.
scalable oversightappellate hierarchyanalogyMost decisions final at the lowest level; review above. Appeals are party-initiated, oversight is sampled.
human-in-the-loopright to human decisionportGDPR art. 22.
corrigibilityrevocability of agencyportThe principal’s unilateral power to amend or terminate the agency.
off-switchpower of terminationportHadfield-Menell et al. (2017) formalize when the agent submits to it.
deceptive alignmentfraudulent concealmentanalogyRequires a mental state that doctrine cannot yet locate.
sandbaggingmisrepresentation in examinationanalogy
treacherous turnsleeper agentanalogyAn espionage concept; doctrine has no counterpart.
refusalconscience clauseanalogyThe common-carrier duty to serve is the more instructive contrast.
guardrailsregulation compiled to architectureportLessig (1999), executed literally.
responsible scaling policyinternal compliance programsameSelf-regulation drafted in anticipation of statute.
deployment gatepermittingport
safety caseburden of proofportThe developer shows safety; the regulator need not show harm.
AI safety levelsbiosafety levelsportRegulatory classification imported with the acronym.
red-team safe harborsecurity research exemptionsameThe DMCA §1201 shape, requested for models.
CSAM reportingmandatory reporting statutesame18 U.S.C. §2258A already governs providers.
long-term benefit trusttrustsameAnthropic’s LTBT is a trust in the ordinary legal sense.
model welfareanimal welfare lawanalogyDuties regarding an entity, possibly duties toward one.

What does not transfer

Law presupposes four things the substrate does not supply.

Are safety and alignment just law and compliance?

Sources