<?xml version='1.0' encoding='utf-8'?>
<!DOCTYPE rfc [
  <!ENTITY nbsp    "&#160;">
  <!ENTITY zwsp   "&#8203;">
  <!ENTITY nbhy   "&#8209;">
  <!ENTITY wj     "&#8288;">
]>
<!-- name="GENERATOR" content="github.com/mmarkdown/mmark Mmark Markdown Processor - mmark.miek.nl" -->
<rfc xmlns:xi="http://www.w3.org/2001/XInclude" version="3" ipr="trust200902" docName="draft-hardt-aauth-budgets-00" submissionType="IETF" category="std" xml:lang="en" indexInclude="true">

<front>
<title abbrev="AAuth-Budgets">AAuth Budgets</title><seriesInfo value="draft-hardt-aauth-budgets-00" stream="IETF" status="standard" name="Internet-Draft"/>
<author initials="D." surname="Hardt" fullname="Dick Hardt"><organization>Hellō</organization><address><postal><street/>
</postal><email>dick.hardt@gmail.com</email>
</address></author><date/>
<area>Security</area>
<workgroup>TBD</workgroup>
<keyword>agent</keyword>
<keyword>authorization</keyword>
<keyword>budget</keyword>
<keyword>metering</keyword>
<keyword>http</keyword>
<keyword>resource</keyword>

<abstract>
<t>This document defines AAuth Budgets, an extension to the AAuth Protocol (<xref target="I-D.hardt-oauth-aauth-protocol"/>) that carries a spending ceiling from a person server to a resource. A budget is a ceiling on what an agent may consume at one resource, denominated in a unit the resource declares, carried as a claim in the auth token, and enforced by the resource. Budgets are structurally parallel to scope: the agent asks, the resource offers, the person server and access server may narrow, and the auth token carries what was granted. The extension adds a <tt>budget</tt> claim to resource tokens and auth tokens, a <tt>budget_consumed</tt> claim reporting what the presented auth token consumed, a <tt>budget_units</tt> field and a <tt>usage_endpoint</tt> to resource metadata, and an <tt>AAuth-Budget</tt> response header reporting what a request cost and what remains.</t>
</abstract>

<note><name>Discussion Venues</name>
<t><em>Note: This section is to be removed before publishing as an RFC.</em></t>
<t>This document is part of the AAuth specification family. Source for this draft and an issue tracker can be found at <eref target="https://github.com/dickhardt/AAuth">https://github.com/dickhardt/AAuth</eref>.</t>
</note>

</front>

<middle>

<section anchor="introduction"><name>Introduction</name>
<t><strong>Status: Exploratory Draft</strong></t>
<t>The AAuth Protocol (<xref target="I-D.hardt-oauth-aauth-protocol"/>) lets a person server (PS) decide whether an agent may access a resource, and lets the resource express what access it offers. Neither party has a way to say <em>how much</em>.</t>
<t>For a resource that meters and charges per call, that omission is the whole authorization decision. An agent harness calling a model inference endpoint on a person's account can spend without bound: the scope <tt>inference.completions</tt> is either granted or not, and once granted it says nothing about whether the agent may consume ten cents or ten thousand dollars of the person's money. The person's only controls are outside the protocol — a provider dashboard, a card limit, a bill that arrives after the fact.</t>
<t>This document defines a <strong>budget</strong>: a ceiling on what an agent may consume at one resource, denominated in a unit the resource declares, carried as a claim in the auth token, and enforced by the resource.</t>
<t>A budget is an authorization, not a hint. The PS has authorized the agent to spend up to a stated amount, and the resource is the party that counts. That distinction determines nearly every design choice in this document, in particular why the balance cannot be reported through <tt>RateLimit</tt> (<xref target="I-D.ietf-httpapi-ratelimit-headers"/>) — see <xref target="why-not-ratelimit"/>.</t>
<t>The granted budget is an allocation, not the person's ceiling. The ceiling is the person server's own state: no claim carries it, and it may not be shared with the agent. The PS sizes each auth token against what the work has cost so far, and the token's expiry or its budget's exhaustion — whichever comes first — brings the agent back for the next allocation. That return is the supervision point, and the PS's options there are the subject of <xref target="ps-token-endpoint"/>. Nobody knows at mission approval what an agent's work will cost; a figure fixed once up front is either too small to finish or too large to be a control <xref target="why-not-the-ceiling"/>.</t>
<t>Metered inference is the initiating use case, and <xref target="inference"/> covers it as a named deployment pattern. The mechanism is general: any resource that meters and charges per call uses it unchanged.</t>
<t>TPX <xref target="TPX"/> profiles the same grant for OAuth 2.0: a person grants a human-driven app a metered inference budget — "a damage cap, not a payment" — from a provider the person chooses and pays. This document is the AAuth counterpart: the same grant, carried to an autonomous agent through the narrowing chain, and generalized beyond inference to any resource that meters. A provider implementing both accepts two authorization envelopes over one meter <xref target="inference"/>.</t>

<section anchor="non-goals"><name>Non-Goals</name>

<ul spacing="compact">
<li><strong>Not a mission aggregate.</strong> A budget covers a single resource in a single unit. It does not express a cross-resource total such as "$5,000 for the Japan trip." Mission-wide totals require aggregation across resources that meter in different units, which this document does not define.</li>
<li><strong>Not a rate limit.</strong> A budget is cumulative consumption, not per-window throughput. <tt>RateLimit</tt> and <tt>RateLimit-Policy</tt> (<xref target="I-D.ietf-httpapi-ratelimit-headers"/>) cover throughput. Both MAY appear on the same response as <tt>AAuth-Budget</tt>, meaning different things.</li>
<li><strong>Not pricing.</strong> The resource prices its own service. A budget bounds spend at whatever prices the resource charges; this document defines no way to express a price.</li>
<li><strong>Not composite.</strong> A budget is one amount in one unit. A resource that meters several quantities at different rates — input tokens, output tokens, cache reads — collapses them to one billing unit, typically currency, before denominating a budget.</li>
<li><strong>Not payment or settlement.</strong> No funds move. <tt>402 Payment Required</tt> and the resource's commercial arrangement with the person are untouched.</li>
<li><strong>Not PS-enforced at request time.</strong> The PS authorizes a number. The resource counts against it. The PS is not in the request path.</li>
<li><strong>Not an OAuth extension.</strong> This document defines claims in AAuth tokens (<tt>aa-resource+jwt</tt>, <tt>aa-auth+jwt</tt>), fields in <tt>aauth-resource.json</tt>, a resource endpoint, an AAuth capability value, and an AAuth response header. It registers nothing in an OAuth registry. The documents surveyed in <xref target="prior-art"/> are cited as prior art and are non-normative.</li>
</ul>
</section>
</section>

<section anchor="conventions-and-definitions"><name>Conventions and Definitions</name>
<t>{::boilerplate bcp14-tagged}</t>
</section>

<section anchor="terminology"><name>Terminology</name>

<ul spacing="compact">
<li><strong>Budget</strong>: A ceiling on what an agent may consume at one resource, expressed as an amount in a unit the resource declares.</li>
<li><strong>Unit</strong>: A resource-declared identifier for what is being metered — a currency code, a token count, a call count.</li>
<li><strong>Scale</strong>: The number of decimal places implied by a budget amount, carried as <tt>decimals</tt>. An amount of <tt>5000000</tt> with <tt>decimals</tt> of <tt>6</tt> is 5.000000 of the unit.</li>
<li><strong>Granted budget</strong>: The <tt>budget</tt> claim of an auth token. The figure the resource enforces against.</li>
<li><strong>Consumption</strong>: What the resource has metered against a granted budget, in the same unit and scale.</li>
<li><strong>Consumption record</strong>: A <tt>{jti, consumed}</tt> pair reporting what the resource metered against one auth token's budget, as of the resource token that carries it. Carried in the <tt>budget_consumed</tt> claim of a resource token.</li>
<li><strong>Usage counters</strong>: Calendar-aligned consumption totals a person server reads at the resource's <tt>usage_endpoint</tt>.</li>
</ul>
</section>

<section anchor="budget-model"><name>Budget Model</name>

<section anchor="budget-is-scope"><name>A Budget Is Structurally a Scope</name>
<t>The AAuth Protocol defines <tt>scope</tt> in three positions with a narrowing rule (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Scopes). A budget occupies the same three positions, plus the PS-to-AS hop in four-party access:</t>
<table>
<thead>
<tr>
<th>Position</th>
<th><tt>scope</tt></th>
<th><tt>budget</tt></th>
</tr>
</thead>

<tbody>
<tr>
<td>Authorization endpoint request</td>
<td>what the agent asks for</td>
<td>what the agent asks for</td>
</tr>

<tr>
<td>Resource token</td>
<td>what the resource will grant</td>
<td>what the resource will grant</td>
</tr>

<tr>
<td>PS-to-AS token request</td>
<td>(n/a)</td>
<td>what the PS will allow</td>
</tr>

<tr>
<td>Auth token</td>
<td>granted, MUST NOT be broader</td>
<td>granted, MUST NOT exceed</td>
</tr>
</tbody>
</table><t>The base protocol's rule that a resource token MUST only include resource scopes the resource has declared in its <tt>scope_descriptions</tt> metadata has a direct parallel here: a resource token MUST only name a unit the resource has declared in <tt>budget_units</tt> <xref target="budget-units"/>.</t>
<t>This extension therefore introduces one new claim shape and no new authorization semantics. Every party that already knows how to narrow a scope knows how to narrow a budget.</t>
</section>

<section anchor="narrowing-chain"><name>The Narrowing Chain</name>
<t>Each stage MUST NOT exceed the previous stage. There is one asymmetry between the amount and the denomination: the resource settles the denomination, and no later party may change it.</t>

<sourcecode type="ascii-art"><![CDATA[Agent            Resource          PS               AS
  |                 |               |                |
  | budget request  |               |                |
  | (OPTIONAL)      |               |                |
  |---------------->|               |                |
  |                 | sets unit and decimals,        |
  |                 | MAY lower amount               |
  |                 |               |                |
  |  resource token (budget)        |                |
  |<----------------|               |                |
  |                 |               |                |
  |  resource token |               |                |
  |-------------------------------->|                |
  |                 |               | MAY lower      |
  |                 |               | amount         |
  |                 |               |                |
  |                 |     budget (four-party only)   |
  |                 |               |--------------->|
  |                 |               |                | MAY lower
  |                 |               |                | amount
  |                 |               |  auth token    |
  |                 |               |<---------------|
  |  auth token (granted budget)    |                |
  |<--------------------------------|                |
]]>
</sourcecode>
<t>{: #fig-narrowing title="Budget narrowing. Only the resource sets the unit and scale."}</t>

<ol>
<li><t><strong>The agent requests.</strong> OPTIONAL. A <tt>budget</tt> object in the authorization endpoint request <xref target="authorization-endpoint"/>. Omitting it means the resource applies its own default. Requesting more than the resource will allow is NOT an error; the resource narrows.</t>
</li>
<li><t><strong>The resource sets the unit and scale, and MAY lower the amount.</strong> The resource is the enforcer and the only party that knows its own pricing, so it settles the denomination. It MAY change <tt>unit</tt> from what the agent requested — for example converting a request denominated in tokens into a currency amount. After this stage, <tt>unit</tt> and <tt>decimals</tt> are fixed for the life of the grant.</t>
</li>
<li><t><strong>The PS MAY lower the amount.</strong> The PS MUST NOT change <tt>unit</tt> or <tt>decimals</tt>. It is applying the person's policy to a figure the resource denominated; a PS that redenominated would be stating a budget in something the resource may not meter in.</t>
</li>
<li><t><strong>The AS MAY lower the amount</strong> (four-party only), for its own credit or risk reasons. The same prohibition on changing <tt>unit</tt> or <tt>decimals</tt> applies.</t>
</li>
</ol>
<t>Because the resource MAY change the unit, an agent MUST NOT assume the granted budget is directly comparable to what it requested. The agent reads what it actually got from the <tt>budget</tt> claim of its auth token.</t>
<t>The resource token carries one figure, not both the agent's request and the resource's own maximum. It is the minimum of the two, exactly as <tt>scope</tt> is. What the agent originally asked for has no bearing on the PS's decision; an agent that wants the PS to know it belongs in <tt>justification</tt> (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Auth Token Endpoint).</t>
</section>

<section anchor="value-representation"><name>Value Representation</name>
<t>A budget amount is a <strong>non-negative integer in a scale the resource declares</strong>. It is not a floating-point number and not a decimal string.</t>
<t>The value of a budget is <tt>amount</tt> divided by 10 raised to the power of <tt>decimals</tt>, in <tt>unit</tt>. An <tt>amount</tt> of <tt>5000000</tt> with <tt>unit</tt> of <tt>USD</tt> and <tt>decimals</tt> of <tt>6</tt> is five US dollars.</t>
<t>The scale is declared per unit, not globally. A resource metering US dollars per inference call wants <tt>decimals</tt> of 6 to represent micro-dollars; a resource metering Japanese yen wants 0, matching the ISO 4217 <xref target="ISO4217"/> minor unit; a resource metering tokens wants 0.</t>
<t><tt>decimals</tt> is the name x402 <xref target="x402"/> and ERC-20 <xref target="ERC20"/> use for this quantity. ISO 4217 and ISO 20022 call it the minor unit. This document does not introduce a third name. See <xref target="why-integer"/> for why the value is an integer and <xref target="why-explicit-scale"/> for why the scale travels with the amount rather than being derived from the unit identifier.</t>
</section>

<section anchor="range-limits"><name>Range Limits</name>
<t>Two carriers bound the representable range:</t>

<ul spacing="compact">
<li>A Structured Field Integer (<xref target="RFC9651"/>, Section 3.3.1) is limited to 15 digits, which bounds <tt>remaining</tt>, <tt>cost</tt>, and <tt>reserved</tt> in the <tt>AAuth-Budget</tt> header <xref target="aauth-budget-header"/>.</li>
<li>A JSON number is exact only to 2^53 (approximately 9.0 x 10^15), which bounds the <tt>amount</tt> member of the <tt>budget</tt> claim.</li>
</ul>
<t>US dollars at <tt>decimals</tt> of 6 therefore top out near 10^9 per budget, which is well beyond any plausible grant. Implementations MUST NOT issue a budget whose <tt>amount</tt> exceeds 999,999,999,999,999 (15 digits), so that the amount and every derived figure remain representable in both carriers.</t>
<t>If assets requiring 18 decimal places come into scope, both carriers break and <tt>amount</tt> would have to become a string, as it is in x402. This document does not define that change.</t>
</section>

<section anchor="budget-object"><name>The Budget Object</name>
<t>One shape appears in every position — the authorization endpoint request, the resource token, the PS-to-AS token request, and the auth token:</t>

<sourcecode type="json"><![CDATA[{
  "amount": 5000000,
  "unit": "USD",
  "decimals": 6
}
]]>
</sourcecode>
<t>Members:</t>

<ul spacing="compact">
<li><strong><tt>amount</tt></strong> (REQUIRED). A non-negative integer, subject to <xref target="range-limits"/>.</li>
<li><strong><tt>unit</tt></strong> (REQUIRED). A string identifying what is metered. For currency, the value SHOULD be an ISO 4217 <xref target="ISO4217"/> alphabetic code. Units are resource-declared; this document establishes no registry of unit values, for the same reason the base protocol establishes no registry of scope values.</li>
<li><strong><tt>decimals</tt></strong> (REQUIRED). A non-negative integer giving the scale of <tt>amount</tt>. When <tt>unit</tt> is an ISO 4217 code, <tt>decimals</tt> is NOT constrained to that currency's minor unit; see <xref target="why-explicit-scale"/>.</li>
</ul>
<t>The object appears alongside <tt>scope</tt> wherever both are present:</t>

<sourcecode type="json"><![CDATA[{
  "scope": "inference.completions",
  "budget": { "amount": 5000000, "unit": "USD", "decimals": 6 }
}
]]>
</sourcecode>
</section>
</section>

<section anchor="budget-units"><name>Resource Metadata Extensions</name>
<t>This document extends the <tt>/.well-known/aauth-resource.json</tt> document defined in AAuth Protocol (<xref target="I-D.hardt-oauth-aauth-protocol"/>) with two fields:</t>

<sourcecode type="json"><![CDATA[{
  "issuer": "https://inference.example",
  "jwks_uri": "https://inference.example/.well-known/jwks.json",
  "access_mode": "auth-token",
  "authorization_endpoint": "https://inference.example/authorize",
  "scope_descriptions": {
    "inference.completions": "Generate completions,
      billed to your account"
  },
  "budget_units": [
    { "unit": "USD", "decimals": 6, "max": 10000000 },
    { "unit": "tokens", "decimals": 0, "max": 5000000 }
  ],
  "usage_endpoint": "https://inference.example/usage",
  "revocation_endpoint": "https://inference.example/revoke"
}
]]>
</sourcecode>
<t><strong><tt>budget_units</tt></strong> (OPTIONAL). An array of objects, each declaring one unit the resource will denominate a budget in. Each object contains:</t>

<ul spacing="compact">
<li><strong><tt>unit</tt></strong> (REQUIRED). The unit identifier.</li>
<li><strong><tt>decimals</tt></strong> (REQUIRED). The scale the resource uses for this unit. A resource MUST use this value in every <tt>budget</tt> it issues for this unit.</li>
<li><strong><tt>max</tt></strong> (RECOMMENDED). A non-negative integer, in this unit's scale, giving the largest amount the resource will accept on a single auth token. It lets an agent on the proactive path request something the resource will honor rather than discovering the ceiling by having its request narrowed.</li>
<li><strong><tt>description</tt></strong> (OPTIONAL). A Markdown string describing what the unit meters, for display at a consent screen. Implementations MUST sanitize the Markdown before rendering to users.</li>
</ul>
<t>A resource that declares <tt>budget_units</tt> MUST NOT issue a resource token whose <tt>budget.unit</tt> is absent from the array, and MUST set <tt>budget.decimals</tt> to the value declared for that unit.</t>
<t><strong><tt>usage_endpoint</tt></strong> (OPTIONAL). The HTTPS URL where a person server or access server queries usage counters <xref target="usage-counters"/>. The URL MUST conform to the requirements the base protocol places on endpoint URLs (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Endpoint and Other URLs).</t>
<t>The example also carries <tt>revocation_endpoint</tt>, which the base protocol recommends for a resource that accepts person tokens, as a metered resource does (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Resource Metadata). Settlement of a revoked token depends on it <xref target="settlement"/>.</t>
<t>A PS or AS that receives a resource token carrying <tt>budget</tt> already fetches <tt>{iss}/.well-known/aauth-resource.json</tt> to discover the resource's JWKS, per the <tt>dwk</tt> claim (<xref target="I-D.hardt-httpbis-signature-key"/>). The unit declarations arrive in a fetch it was already making, so the cross-check in <xref target="errors"/> costs no extra round trip.</t>
</section>

<section anchor="authorization-endpoint"><name>Authorization Endpoint Extensions</name>
<t>An agent MAY include a <tt>budget</tt> object <xref target="budget-object"/> in the authorization endpoint request body alongside <tt>scope</tt>:</t>

<sourcecode type="http"><![CDATA[POST /authorize HTTP/1.1
Host: inference.example
Content-Type: application/json
AAuth-Capabilities: interaction, budget
Signature-Input: sig=("@method" "@authority"
    "@path" "signature-key");created=1754611200
Signature: sig=:...signature bytes...:
Signature-Key: sig=jwt;jwt="eyJhbGciOiJFZDI1NTE5IiwidHlwIjoiYWEtcGVyc29uK2p3dCIsImtpZCI6InBzLWtleS0xIn0..."

{
  "scope": "inference.completions",
  "budget": { "amount": 10000000, "unit": "USD", "decimals": 6 }
}
]]>
</sourcecode>
<t>The agent presents a person token at the authorization endpoint, as the base protocol requires (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Authorization Endpoint).</t>
<t><strong><tt>budget</tt></strong> (OPTIONAL). The ceiling the agent is requesting. All three members of the budget object are REQUIRED when <tt>budget</tt> is present.</t>
<t>The resource MUST NOT reject the request because <tt>budget.amount</tt> exceeds what it will grant; it narrows instead <xref target="narrowing-chain"/>. The resource MAY reject a <tt>budget</tt> whose <tt>unit</tt> it has not declared <xref target="errors"/>.</t>
<t>When the agent obtains its resource token from a <tt>401</tt> challenge rather than the authorization endpoint (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Auth Token Required), it has not stated a budget and the resource sizes the resource token on its own. This is why <tt>max</tt> in <tt>budget_units</tt> matters: it is what makes the proactive path useful on a first attempt.</t>
<t>How an agent knows an operation is metered before its first call is answered by R3 (<xref target="I-D.hardt-aauth-r3"/>): a resource MAY annotate individual operations in its vocabulary with a budget annotation, and an agent that reads one knows to include <tt>budget</tt> in this request. The annotation states the fact of metering; <tt>budget_units</tt> states the units and ceilings.</t>
</section>

<section anchor="resource-token"><name>Resource Token Extensions</name>
<t>This document extends the resource token (a JWT with <tt>typ: aa-resource+jwt</tt>) with two optional claims.</t>

<ul spacing="compact">
<li><strong><tt>budget</tt></strong> (OPTIONAL): A budget object <xref target="budget-object"/>. The ceiling the resource is willing to have granted. Subject to the rules in <xref target="budget-units"/>.</li>
<li><strong><tt>budget_consumed</tt></strong> (OPTIONAL): A consumption record <xref target="budget-consumed"/> for the auth token presented on the request this resource token answers — what it has consumed to date. Not a grant.</li>
</ul>

<sourcecode type="json"><![CDATA[{
  "iss": "https://inference.example",
  "dwk": "aauth-resource.json",
  "aud": "https://ps.example",
  "jti": "rt-4c81fa",
  "ps": "https://ps.example",
  "sub": "8f14e45fceea167a5a36dedd4bea2543",
  "presented_jti": "at-71b9d0",
  "agent_jkt": "NzbLsXh8uDCcd-6MNwXF4W_7noWXFZAfHkxZsRGC9Xs",
  "tenant": "corp",
  "mission_s256": "dBjftJeZ4CVP-mB92K27uhbUJU1p1r_wW1gFWFOEjXk",
  "scope": "inference.completions",
  "budget": { "amount": 5000000, "unit": "USD", "decimals": 6 },
  "budget_consumed": { "jti": "at-71b9d0", "consumed": 2000000 },
  "iat": 1754612400,
  "exp": 1754612700
}
]]>
</sourcecode>
<t>This example is the resource token a resource issues on a challenge to a request presenting the auth token of <xref target="auth-token"/> once that token's grant is spent, so <tt>presented_jti</tt> and the record both name that token.</t>

<section anchor="budget-consumed"><name>The Consumption Record</name>
<t><tt>budget_consumed</tt> is a single consumption record with two members, both REQUIRED:</t>

<ul spacing="compact">
<li><strong><tt>jti</tt></strong>: The <tt>jti</tt> claim of the auth token presented on the request this resource token answers. This is the value the resource token carries as <tt>presented_jti</tt> (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Resource Token Structure).</li>
<li><strong><tt>consumed</tt></strong>: What the resource has metered against that token's budget, as of this resource token's <tt>iat</tt>: a non-negative integer in the <tt>unit</tt> and <tt>decimals</tt> of the <tt>budget</tt> claim in the same resource token.</li>
</ul>
<t>A resource MUST NOT include <tt>budget_consumed</tt> unless <tt>budget</tt> is present in the same token, and MUST omit it rather than state it in another scale — a record carried in a stale scale is the thousandfold error <xref target="errors"/> in miniature.</t>
<t>The record exists for one figure the issuer needs and cannot compute. <tt>budget-exhausted</tt> implies the whole grant was spent, but under <tt>insufficient-budget</tt> <xref target="exhaustion"/>, and on a token that expired with budget left, the token's spend to date is a number only the resource holds — and without it the issuer accounts for the allocation as fully consumed <xref target="unreported-allocations"/> when most of it may remain. A resource issuing a resource token on a challenge to a request that carried an auth token SHOULD include the record. A resource issuing one on a person token, at its authorization endpoint or on the first challenge of a grant, has no auth token to report on and omits the claim. In four-party access the resource token travels to the AS inside the PS's token request, so the record reaches both issuers without a usage query.</t>
<t>The record is as of the resource token's <tt>iat</tt>. It is a snapshot: the token it names is still valid, since a resource issues a resource token only on a request carrying a valid person token or auth token, and the token may spend more before it expires. A later record for the same <tt>jti</tt> supersedes an earlier one. A token's final figure is settled by the usage reading <xref target="settlement"/>.</t>
<t>A record is deliberately two members and no more. The <tt>jti</tt> it names is a token the PS issued — or, in four-party access, relayed from the AS — so the issuer already holds the granted amount, the mission, the scope, and the issuance time, and joins them from its own ledger. Carrying those values again would duplicate what the issuer knows and put more of the person's financial detail into a token the agent also reads. What the issuer cannot know, and what the record supplies, is what the resource actually metered.</t>
<t>The record rides in the resource token because it already travels resource → agent → PS at exactly the moment the PS re-decides: no extra round trip, resource-signed, and interpretable without a metadata fetch. It names the presented token and nothing else. The spend under a person's other tokens is served at the usage endpoint <xref target="usage-counters"/>, which answers that question better — per key, per mission, over calendar periods — on a channel the agent is not on. An earlier revision carried up to twenty records; <xref target="why-one-record"/> says why one is enough.</t>
</section>
</section>

<section anchor="ps-token-endpoint"><name>Auth Token Endpoint Extensions</name>
<t>No new request parameter is defined for the agent's request to the PS's <tt>auth_token_endpoint</tt>. The budget reaches the PS inside the resource token.</t>
<t>An agent seeking a larger budget obtains a fresh resource token from the resource — stating the larger figure in the authorization endpoint request <xref target="authorization-endpoint"/>, or being handed one on a <tt>401</tt> <xref target="exhaustion"/> — and SHOULD explain the need in the <tt>justification</tt> parameter of the auth token request.</t>
<t>When the PS issues the auth token itself (three-party), it applies the person's policy and issues per <xref target="auth-token"/>. When it federates (four-party), it proceeds per <xref target="as-token-endpoint"/>.</t>

<section anchor="ps-decision"><name>What the PS Is Deciding</name>
<t>The resource token's <tt>budget</tt> states what the resource will allow. It is an offer, not a request the PS is obliged to answer in full.</t>
<t>Against that offer the PS holds a ceiling for the person at this resource — a standing limit, a mission's stated intent, an organizational policy, or a figure the person supplied when asked. The ceiling is PS state. This document defines no wire format for it, no claim that carries it, and no way for the agent to read it. What the PS issues is an allocation drawn against it.</t>
<t>Sizing the allocation is where the PS's supervision happens. A PS that issues the resource's full offer every time has authorized the resource's maximum and learns nothing until the money is gone. A PS that issues a fraction sees the agent again when that fraction is spent, with the token's consumption record in hand, and decides then whether the work is going as the person expected.</t>
<t>The interval is not fixed by the clock. An auth token expires within an hour, and its budget is exhausted after however much work it took to spend — whichever comes first returns the agent to the PS. A mission running cheaply reports on the hour; one running expensively reports in minutes. The PS sets that frequency by sizing the allocation, and no party configures it <xref target="token-scope"/>.</t>
</section>

<section anchor="ps-inputs"><name>What the PS Reads</name>
<t>Four inputs are available at the moment of the decision, and a PS applying the person's policy SHOULD use all of them:</t>

<ul spacing="compact">
<li><strong><tt>budget_consumed</tt></strong> <xref target="budget-consumed"/> in the resource token the agent just presented: what the token it was presenting has cost so far, resource-signed, arriving at no round-trip cost.</li>
<li><strong>Usage counters</strong> <xref target="usage-counters"/> at the resource's <tt>usage_endpoint</tt>: totals over calendar periods, and for a mission query the mission's total to date — the figures that cover the stretch when the agent was not talking to the PS.</li>
<li><strong>The mission log</strong> (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Mission Log): every prior token request, justification, and clarification in this mission, which is what makes "faster than expected" a judgement the PS can actually make.</li>
<li><strong>The <tt>justification</tt></strong> parameter of this request: why the agent says it needs more.</li>
</ul>
<t>The first two are the spend; the second two are the context. A budget escalation is not interpretable without both.</t>
</section>

<section anchor="ps-responses"><name>How the PS Responds</name>
<t>Six responses are available. None is new to this document; the base protocol defines each, and this section states which apply to a budget decision.</t>
<table>
<thead>
<tr>
<th>Response</th>
<th>Mechanism</th>
</tr>
</thead>

<tbody>
<tr>
<td>Grant the offer</td>
<td>Issue an auth token with <tt>budget</tt> equal to the resource token's <xref target="auth-token"/></td>
</tr>

<tr>
<td>Grant less</td>
<td>Issue a lower <tt>amount</tt> <xref target="narrowing-chain"/></td>
</tr>

<tr>
<td>Ask the agent</td>
<td><tt>202</tt> with <tt>requirement=clarification</tt></td>
</tr>

<tr>
<td>Ask the person</td>
<td><tt>202</tt> with <tt>requirement=interaction</tt></td>
</tr>

<tr>
<td>Decline with a figure</td>
<td>Error response carrying <tt>suggested_budget</tt> <xref target="declining"/></td>
</tr>

<tr>
<td>End the work</td>
<td>Terminate the mission</td>
</tr>
</tbody>
</table><t>Granting less needs no signalling: the <tt>amount</tt> in the issued claim is the answer, and the agent reads it from the token it received <xref target="narrowing-chain"/>.</t>
<t><strong>Clarification is the response for an escalation the PS is not ready to refuse or approve.</strong> A PS that sees consumption running ahead of what the mission implies MAY return <tt>202</tt> with <tt>requirement=clarification</tt> (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Clarification Required), putting a question to the agent before deciding. This is the channel that lets the PS tell an agent it is overspending, which narrowing alone cannot do — a smaller <tt>amount</tt> is silent, and the agent cannot distinguish a PS applying pressure from a resource lowering its own offer.</t>

<sourcecode type="http"><![CDATA[HTTP/1.1 202 Accepted
Location: /pending/abc123
Retry-After: 0
Cache-Control: no-store
AAuth-Requirement: requirement=clarification
Content-Type: application/json

{
  "status": "pending",
  "clarification": "This mission has spent $18 of an
    expected $25 and has not booked anything yet. What
    is the remaining $12 for?",
  "timeout": 120
}
]]>
</sourcecode>
<t>The agent's three replies are already defined and all three are useful here: a <tt>clarification_response</tt> explaining the spend, an <tt>updated_request</tt> carrying a fresh resource token for a smaller figure with the <tt>presented_token</tt> it used to obtain it (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Updated Request), or a <tt>DELETE</tt> withdrawing the request. The base protocol's limit on clarification rounds applies (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Clarification Limits). An agent that did not declare the <tt>clarification</tt> capability cannot be asked, and the PS decides without it.</t>
<t>Asking the person is the same mechanism one step further out, and is the right response when the answer is the person's rather than the agent's — a ceiling raise rather than an allocation.</t>
<t>A PS that puts a budget to the person for consent MUST present the amount as a human-readable figure in its unit — "$5.00", not <tt>{5000000, USD, 6}</tt> — visually distinct from any resource-supplied description, which is Markdown and MUST be sanitized before rendering <xref target="budget-units"/>. The amount is the decision the person is making.</t>
<t>Ending the work is the response to an agent whose spending the PS cannot account for. Terminating a mission is not a budget mechanism and this document defines nothing about it; it is named here because a budget escalation is one of the few signals that reliably surfaces an agent behaving unlike its mission.</t>
</section>
</section>

<section anchor="as-token-endpoint"><name>PS-to-AS Token Request Extensions</name>
<t>This document extends the PS-to-AS token request (<xref target="I-D.hardt-oauth-aauth-protocol"/>, PS-to-AS Token Request) with one parameter, used in four-party access only.</t>
<t><strong><tt>budget</tt></strong> (OPTIONAL): A budget object <xref target="budget-object"/> carrying the ceiling the PS will allow.</t>

<sourcecode type="http"><![CDATA[POST /token HTTP/1.1
Host: as.inference.example
Content-Type: application/json
Content-Digest: sha-256=:...:
Signature-Input: sig=("@method" "@authority" "@path"
    "content-type" "content-digest"
    "signature-key");created=1754611200
Signature: sig=:...signature bytes...:
Signature-Key: sig=jwks_uri;id="https://ps.example";
    dwk="aauth-person.json";kid="key-1"

{
  "resource_token": "eyJhbGc...",
  "agent_token": "eyJhbGc...",
  "presented_token": "eyJhbGc...",
  "budget": { "amount": 2000000, "unit": "USD", "decimals": 6 }
}
]]>
</sourcecode>
<t>The PS MUST copy <tt>unit</tt> and <tt>decimals</tt> from the resource token's <tt>budget</tt> claim unchanged, and MUST NOT set <tt>amount</tt> higher than the resource token's <tt>budget.amount</tt>. The AS MUST NOT issue a <tt>budget</tt> claim exceeding this parameter, and MAY lower it further.</t>
<t>When the AS returns an auth token, the PS MUST verify, along with the checks of the base protocol's Auth Token Delivery, that the token carries a <tt>budget</tt> claim when the PS sent a <tt>budget</tt> parameter, that the claim has the <tt>unit</tt> and <tt>decimals</tt> of that parameter and an <tt>amount</tt> no higher, and that the token carries no <tt>budget</tt> claim when the PS sent no <tt>budget</tt> parameter. A token without a <tt>budget</tt> claim would leave the resource to apply its own default, which can exceed what the PS allowed. A token that fails is an auth token that fails delivery verification, and the PS answers the agent <tt>as_unreachable</tt> (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Auth Token Delivery). This is the budget counterpart of the base protocol's check that the AS's <tt>scope</tt> is no broader than the resource token's.</t>
<t>When the resource token carries <tt>budget</tt> and the PS omits this parameter, the AS MUST NOT issue a <tt>budget</tt> claim. A PS that grants the resource's full offer says so by echoing the resource token's <tt>budget</tt>; omission is what a PS that does not implement this extension sends, and it MUST NOT be read as a grant. The auth token then carries no allocation, and the resource applies its own default to it <xref target="auth-token"/>. Reading omission as the full offer would turn a PS's non-participation into the maximum grant, the opposite of what ignoring an unrecognized claim is meant to do.</t>
</section>

<section anchor="auth-token"><name>Auth Token Extensions</name>
<t>This document extends the auth token (a JWT with <tt>typ: aa-auth+jwt</tt>) with one optional claim.</t>
<t><strong><tt>budget</tt></strong> (OPTIONAL): A budget object <xref target="budget-object"/>. This is the granted budget and is authoritative.</t>

<sourcecode type="json"><![CDATA[{
  "iss": "https://ps.example",
  "dwk": "aauth-person.json",
  "aud": "https://inference.example",
  "jti": "at-71b9d0",
  "ps": "https://ps.example",
  "sub": "8f14e45fceea167a5a36dedd4bea2543",
  "cnf": { "jwk": { "kty": "OKP", "crv": "Ed25519",
                    "x": "NzbLsXh8uDCcd...", "alg": "Ed25519" } },
  "tenant": "corp",
  "mission_s256": "dBjftJeZ4CVP-mB92K27uhbUJU1p1r_wW1gFWFOEjXk",
  "scope": "inference.completions",
  "budget": { "amount": 2000000, "unit": "USD", "decimals": 6 },
  "iat": 1754611200,
  "exp": 1754614800
}
]]>
</sourcecode>
<t>Issuer rules:</t>

<ul spacing="compact">
<li>The issuer MUST copy <tt>unit</tt> and <tt>decimals</tt> from the resource token's <tt>budget</tt> claim unchanged.</li>
<li>The issuer MUST NOT set <tt>amount</tt> higher than the resource token's <tt>budget.amount</tt>, or, in four-party access, higher than the <tt>budget</tt> parameter of the PS-to-AS token request <xref target="as-token-endpoint"/>.</li>
<li>An issuer MUST NOT include <tt>budget</tt> in an auth token when the resource token carried no <tt>budget</tt> claim. The resource, not the PS or AS, denominates.</li>
</ul>
<t>Resource rules:</t>

<ul spacing="compact">
<li>A resource that issued a <tt>budget</tt> in a resource token and receives an auth token without a <tt>budget</tt> claim MUST treat the request as carrying no budget authorization under this extension and apply its own default.</li>
<li>A resource MUST enforce against the auth token's <tt>unit</tt> and <tt>decimals</tt>, not against its own current <tt>budget_units</tt> metadata. Auth tokens live up to an hour and metadata can change within that hour; see <xref target="why-signed-decimals"/>.</li>
</ul>
</section>

<section anchor="aauth-budget-header"><name>AAuth-Budget Response Header</name>
<t><tt>AAuth-Budget</tt> is a response header carrying the remaining balance of the granted budget. It is a Dictionary (<xref target="RFC9651"/>, Section 3.2), matching <tt>AAuth-Requirement</tt>.</t>

<sourcecode type="http"><![CDATA[AAuth-Budget: cost=221200, remaining=1568800,
    unit="USD", decimals=6
]]>
</sourcecode>
<t>Members:</t>

<ul spacing="compact">
<li><strong><tt>remaining</tt></strong> (REQUIRED): A non-negative Integer, in the granted scale, giving what is left of the budget on this auth token, net of reservations for requests in flight <xref target="overshoot"/>. It is a floor — committed consumption will not exceed the grant — though the figure may lag metering. Exhaustion is signaled by the <tt>requirement=auth-token</tt> challenge, as a <tt>401</tt> or a <tt>202</tt> <xref target="exhaustion"/>, for which the agent stays prepared regardless.</li>
<li><strong><tt>cost</tt></strong> (OPTIONAL): A non-negative Integer, in the granted scale, giving what <strong>this request</strong> cost. A resource sends it in the header when it knows the figure as it writes the response, in a trailer when it learns the figure after <xref target="streaming"/>, and not at all when it will not learn it in time to do either. The third case is bounded by <xref target="cost-omitted"/>.</li>
<li><strong><tt>reserved</tt></strong> (OPTIONAL): A non-negative Integer, in the granted scale, giving what the resource has held against the grant for this request and not yet committed <xref target="overshoot"/>. Meaningful only where <tt>cost</tt> is not yet known, so in practice it accompanies a streamed response. It is a statement about this request, not a running total, and is never revised. REQUIRED where <tt>cost</tt> is omitted <xref target="cost-omitted"/>.</li>
<li><strong><tt>required</tt></strong> (OPTIONAL): A non-negative Integer, in the granted scale, giving the maximum cost the resource computed for a request it refused under <tt>reason=insufficient-budget</tt> <xref target="reason-parameter"/>. Sent only with that refusal, where it is RECOMMENDED. It is what the request needed, not what the resource is asking the PS to grant next; see <xref target="required-member"/>.</li>
<li><strong><tt>unit</tt></strong> (OPTIONAL): A String naming the unit. <strong><tt>decimals</tt></strong> (OPTIONAL): an Integer giving its scale. Both are informational, and they are a pair: a sender MUST include both or neither. They exist for readers that never parse a JWT — proxies, logs, dashboards. Those are the same readers that would misinterpret an amount carrying a unit with no scale, by a factor of 10^decimals.</li>
</ul>
<t>Recipients MUST ignore members they do not recognize.</t>
<t>The header carries no <tt>granted</tt> member, no cumulative consumption figure, and no token reference. The agent holds the auth token it signed the request with and reads <tt>granted</tt> from there; what it needs per response is what this call cost and what is left, which is what the field carries. Cumulative consumption is an issuer-facing figure, reported in the resource token <xref target="budget-consumed"/> and at the usage endpoint <xref target="usage-counters"/>; see <xref target="why-no-cumulative"/> for why it is not also reported to the agent.</t>
<t>The scope of the reported figures is this auth token's budget, because the budget expires with the token <xref target="enforcement"/>.</t>

<section anchor="header-rules"><name>Sending Rules</name>
<t>A resource that granted a budget SHOULD include <tt>AAuth-Budget</tt> on every response to a request bearing that auth token — success, error, and the challenge, <tt>401</tt> or <tt>202</tt> <xref target="exhaustion-deferred"/>, where it reads <tt>remaining=0</tt> beside the <tt>AAuth-Requirement</tt> header <xref target="exhaustion"/>.</t>
<t>This is SHOULD rather than MUST because the failing layer may sit below the metering layer: a gateway timeout, a crashed worker, or a fault in metering itself produces a response no budget figure can ride on. A resource MUST NOT omit the field for any other reason. An agent that misses the field learns the balance from its next response, and until then applies <xref target="ambiguous-failure"/>.</t>

<sourcecode type="http"><![CDATA[HTTP/1.1 401 Unauthorized
AAuth-Requirement: requirement=auth-token;
    resource-token="eyJ..."; reason=budget-exhausted
AAuth-Budget: remaining=0, unit="USD", decimals=6
]]>
</sourcecode>
<t>That the field appears on every response regardless of status is the point of putting it in a header. The agent reads the same field whether the body is JSON, a server-sent event stream, a streamed completion, or a problem document <xref target="RFC9457"/>.</t>
<t>Intermediaries MUST NOT add, alter, or remove <tt>AAuth-Budget</tt>. The field reports the state of an authorization the resource issued; an intermediary rewriting it is asserting authorization state it does not hold. This is the inverse of the <tt>RateLimit</tt> rule permitting intermediaries to tighten values (<xref target="I-D.ietf-httpapi-ratelimit-headers"/>) — see <xref target="why-not-ratelimit"/>.</t>
<t>Recipients MUST ignore <tt>AAuth-Budget</tt> on a response served from cache with a positive <tt>current_age</tt> (<xref target="RFC9111"/>, Section 4.2.3). This is the one <tt>RateLimit</tt> rule that carries over unchanged.</t>
</section>

<section anchor="unit-conflict"><name>Denomination Conflict</name>
<t>The auth token's <tt>budget</tt> claim is authoritative. A recipient MUST NOT act on a header <tt>unit</tt> or <tt>decimals</tt> that disagrees with the <tt>budget</tt> claim in the auth token it presented, and SHOULD treat the whole field as unreliable for that response.</t>
<t>Including the pair makes the field self-describing for proxies and logs that never parse a JWT, at the cost of duplicating signed values; the conflict rule is that cost made explicit.</t>
</section>
</section>

<section anchor="streaming"><name>Streaming</name>
<t>Response headers are written before the body, and a streamed response's actual cost is known only when the stream ends. The resource therefore cannot state <tt>cost</tt> in the header. What it can state is what it has held: it reserved before serving <xref target="overshoot"/>, and <tt>remaining</tt> is already net of that reservation.</t>
<t>A resource serving a streamed response SHOULD send <tt>reserved</tt> in the header and <tt>cost</tt> in a trailer. A resource that cannot send trailers sends <tt>reserved</tt> alone <xref target="cost-omitted"/>.</t>

<sourcecode type="http"><![CDATA[HTTP/1.1 200 OK
Content-Type: text/event-stream
Trailer: AAuth-Budget
AAuth-Budget: remaining=1568800, reserved=431200,
    unit="USD", decimals=6

   ...stream...

AAuth-Budget: cost=221200
]]>
</sourcecode>
<t>The agent computes the balance after the request as <tt>remaining + reserved - cost</tt>. Here that is 1,778,800: the 431,200 held was not all spent, and the unspent 210,000 returns to the grant.</t>

<section anchor="cost-omitted"><name>When <tt>cost</tt> Is Omitted</name>
<t>Not every resource can send a trailer. Trailers exist only on a chunked or HTTP/2-and-later response, and several widely deployed server runtimes provide no way to emit one at all. A resource in that position knows the cost of a streamed response only after its last opportunity to report it.</t>
<t>Such a resource omits <tt>cost</tt> and MUST send <tt>reserved</tt> in the header. The agent recovers the figure from the following response:</t>
<t><tt>cost</tt> = previous <tt>remaining</tt> + <tt>reserved</tt> - current <tt>remaining</tt></t>
<t>The subtraction works because <tt>remaining</tt> is already net of reservations <xref target="aauth-budget-header"/>: the earlier figure is net of the hold, the later one reflects the commit and the release of the unspent remainder. <tt>reserved</tt> is the term that connects them, which is why it stops being optional here.</t>
<t>It recovers one request's cost only where requests on that token are serial. An agent with several requests in flight on one token recovers the net of everything that settled between the two responses, not the cost of any single one, because every concurrent request moves the same <tt>remaining</tt>. An agent that wants per-request figures from a resource that omits <tt>cost</tt> serializes its requests on that token; an agent that only needs the balance does not have to.</t>
<t>Until that next response arrives the agent applies <xref target="ambiguous-failure"/> and treats the request as having cost the full <tt>reserved</tt> amount. That is the conservative direction, and it is the same rule the agent already applies to a response it never received.</t>
<t>A resource MUST NOT omit both <tt>cost</tt> and <tt>reserved</tt>. That combination reports that a metered request happened and gives the agent no figure for it, neither exact nor conservative.</t>
</section>

<section anchor="trailer-rules"><name>Trailer Rules</name>
<t>A resource MAY send <tt>AAuth-Budget</tt> as a trailer field, subject to three rules:</t>

<ol spacing="compact">
<li>The response MUST list <tt>AAuth-Budget</tt> in a <tt>Trailer</tt> header field (<xref target="RFC9110"/>, Section 6.6.1).</li>
<li>A trailer instance MUST NOT restate a member the header instance carried. It carries <tt>cost</tt> and nothing else.</li>
<li>A recipient MUST NOT treat the trailer as necessary. A response that never delivers one is complete.</li>
</ol>
<t>Rule 2 is what makes the field safe under either way a recipient handles trailers. A recipient that discards trailers keeps the header's <tt>remaining</tt>, which is a floor and therefore correct if conservative. A recipient that merges trailer fields into the header set produces a single Dictionary whose keys do not collide, because no key appears twice. Restating a member is the case that would break: <tt>AAuth-Budget</tt> is a Dictionary, duplicate keys resolve last-wins (<xref target="RFC9651"/>, Section 3.2), and which value won would then depend on whether the recipient merged — a difference no sender can observe or control.</t>
<t>Rule 3 follows from trailers being droppable in transit (<xref target="RFC9110"/>, Section 6.5.2) and from trailers existing only on a chunked or HTTP/2-and-later response. The figure a trailer would have carried is also in the next response's <tt>remaining</tt>, so nothing is lost that is not recovered on the following call.</t>
</section>

<section anchor="ambiguous-failure"><name>Ambiguous Failures</name>
<t>A request whose response never arrives — a dropped connection, an aborted stream — leaves the agent unable to say whether it was metered. No trailer arrives on an aborted stream, and the header that would have carried <tt>cost</tt> is on the response that was lost.</t>
<t>The agent MUST assume the request cost as much as the resource had held for it: <tt>reserved</tt> where it saw one, and otherwise the request's maximum cost as the resource would have bounded it <xref target="overshoot"/>. It carries that assumption until a later response's <tt>remaining</tt> supersedes it, and it SHOULD NOT retry a chargeable request before then.</t>
<t>Assuming the maximum is the conservative direction: an agent that under-assumes plans spending it does not have and discovers the shortfall as a <tt>401</tt> <xref target="exhaustion"/>.</t>
<t>Inference APIs commonly emit final usage in the stream's terminal event. That is application-layer and does not provide the application independence this header exists for, so a resource that emits it and can also send a trailer SHOULD send both. What it does mean is that the exact number exists when the stream ends, and a resource whose runtime offers no trailer has it and no protocol carrier for it until the next response <xref target="cost-omitted"/>.</t>
</section>
</section>

<section anchor="exhaustion"><name>Budget Exhaustion</name>
<t><strong>Auth token expired.</strong> The budget expires with the token. The base protocol has the agent refresh an auth token when fewer than five minutes remain and not present it inside that margin (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Expiry and the Refresh Margin). An expired token presented anyway is answered <tt>401</tt> with <tt>Signature-Error: error=expired_jwt</tt>, and the resource issues no resource token. Either way the agent re-authorizes with a person token at the authorization endpoint (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Re-authorization). No consumption record rides on that path, and an allocation that expired unreported settles from the usage reading <xref target="settlement"/>.</t>
<t><strong>Budget exhausted, token still valid.</strong> <tt>401</tt> with <tt>AAuth-Requirement: requirement=auth-token; resource-token="..."</tt>. The base protocol already permits a resource to return <tt>requirement=auth-token</tt> with a new resource token to a request that already carries an auth token, when the request needs more authorization than the token provides, and requires agents to be prepared for this step-up at any time (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Auth Token Required). Budget exhaustion is that case, and the agent's action is the same as for any step-up: take the fresh resource token to its PS, with the exhausted auth token as <tt>presented_token</tt>. The resource token's consumption record <xref target="budget-consumed"/> is the exhausted token's spend to date. The auth token issued against it expires no later than the exhausted one (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Auth Token Structure), so a step-up renews the allocation but not the token's lifetime. A full-length token comes from re-authorizing with a fresh person token.</t>
<t><strong>Request exceeds the remainder.</strong> The budget has remainder, but this request's maximum cost exceeds it <xref target="overshoot"/>. The same <tt>401</tt> challenge, with <tt>reason=insufficient-budget</tt>. The agent has a second move here that exhaustion does not offer: lower the request's bound to fit the <tt>remaining</tt> reported beside the challenge, and retry on the token it already holds. The <tt>required</tt> member <xref target="required-member"/> is what makes that move a calculation rather than a search.</t>
<t>A request refused under this section MUST NOT draw down the budget or appear in the record and counters. The resource declined to serve it; metering the refusal would make exhaustion self-perpetuating. A refusal therefore carries no <tt>cost</tt>; <tt>required</tt> <xref target="required-member"/> is the figure a refusal reports.</t>
<t>The fresh resource token SHOULD carry the presented token's consumption record <xref target="budget-consumed"/>, its spend to date — the context for deciding whether to authorize more, and the figure that tells the issuer how much of the refused allocation was actually consumed.</t>

<section anchor="reason-parameter"><name>The <tt>reason</tt> Parameter</name>
<t>A resource challenging because the budget is exhausted rather than because the token expired SHOULD include a <tt>reason</tt> parameter on the <tt>requirement</tt> member:</t>

<sourcecode type="http"><![CDATA[HTTP/1.1 401 Unauthorized
AAuth-Requirement: requirement=auth-token;
    resource-token="eyJ..."; reason=budget-exhausted
AAuth-Budget: remaining=0,
    unit="USD", decimals=6
]]>
</sourcecode>
<t><tt>reason</tt> is a Token. This document defines two values:</t>

<ul spacing="compact">
<li><strong><tt>budget-exhausted</tt></strong>: The granted budget is spent.</li>
<li><strong><tt>insufficient-budget</tt></strong>: The budget has remainder, but this request's maximum cost exceeds it <xref target="overshoot"/>.</li>
</ul>
<t>For either value, the enclosed resource token MAY carry a <tt>budget</tt> sized for what the resource would need to see granted. The denial is itself the re-authorization offer.</t>
</section>

<section anchor="required-member"><name>Refusing a Request That Does Not Fit</name>
<t>A resource refusing under <tt>insufficient-budget</tt> has computed the request's maximum cost — <xref target="overshoot"/> requires it to, before serving — and SHOULD report that figure as the <tt>required</tt> member of <tt>AAuth-Budget</tt>:</t>

<sourcecode type="http"><![CDATA[HTTP/1.1 401 Unauthorized
AAuth-Requirement: requirement=auth-token;
    resource-token="eyJ..."; reason=insufficient-budget
AAuth-Budget: remaining=150000, required=400000,
    unit="USD", decimals=6
]]>
</sourcecode>
<t>The agent now knows both halves of the refusal: it has 0.15, and the request needed 0.40. Without <tt>required</tt> it knows only the first, and the move <xref target="exhaustion"/> offers it — lower the request's bound and retry on the token it already holds — becomes a search. It cannot compute the figure itself, because the bound is the resource's own calculation against its own pricing, and no part of this specification requires a resource to publish what an operation costs.</t>
<t><tt>required</tt> is not the same figure as the <tt>budget</tt> claim of the enclosed resource token, and the two SHOULD differ. The resource token's <tt>budget</tt> is a re-authorization offer addressed to the PS, and a resource sizing it for exactly the refused request hands back a grant good for one call. <tt>required</tt> is a fact about the request that was refused, addressed to the agent. The header carries it because the agent is the party that acts on it, and because the retry path it enables does not involve the person server at all.</t>
<t>A resource MAY refuse without <tt>required</tt> — where the operation has no cost bound it is willing to state, or where stating it would disclose pricing the resource does not publish. The agent then falls back to <tt>remaining</tt> alone.</t>
<t>No new <tt>requirement</tt> value is minted. The base protocol says an agent that does not recognize a <tt>requirement</tt> value MUST NOT treat the response as satisfiable and surfaces it as an error, while recipients MUST ignore unknown <em>parameters</em> on the <tt>requirement</tt> member (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Requirement Responses). A new value would hard-fail every budget-unaware agent on a condition that plain <tt>auth-token</tt> resolves correctly. That asymmetry — unknown values fail, unknown parameters are ignored — is why this extension extends by parameter.</t>
<t>An agent that understands the <tt>reason</tt> values knows what to do beyond re-authorizing: for <tt>budget-exhausted</tt>, request a larger budget and say why in the <tt>justification</tt> it sends to its PS; for <tt>insufficient-budget</tt>, either that, or shrink the request and retry without involving the PS at all. An agent that understands neither ignores the parameter and re-authorizes, which is always correct.</t>
</section>

<section anchor="exhaustion-deferred"><name>Deferred Delivery</name>
<t>A resource MAY deliver either refusal as a <tt>202</tt> deferred response rather than a <tt>401</tt>, holding the invocation (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Deferred Delivery). The <tt>AAuth-Requirement</tt> header is the same, <tt>reason</tt> included, and so is the enclosed resource token with its consumption record <xref target="budget-consumed"/>. What differs is what becomes of the request. Under the <tt>401</tt> the resource holds nothing, and the agent sends the request again. Under the <tt>202</tt> the resource holds the invocation, and the agent does not send it again: it obtains an auth token and polls the pending URL, and the resource executes the held invocation on the first poll that presents a valid auth token.</t>

<sourcecode type="http"><![CDATA[HTTP/1.1 202 Accepted
Location: /pending/f7a3b9c
Retry-After: 5
Cache-Control: no-store
AAuth-Requirement: requirement=auth-token;
    resource-token="eyJ..."; reason=budget-exhausted
AAuth-Budget: remaining=0, unit="USD", decimals=6
Content-Type: application/json

{
  "status": "pending"
}
]]>
</sourcecode>
<t>A resource holding an invocation under this section follows four rules:</t>

<ol spacing="compact">
<li><strong>Nothing is reserved while it is held.</strong> For the held invocation, the resource MUST NOT reserve against or meter to the token presented on the original request. For that token the <tt>202</tt> is a refusal like the <tt>401</tt>, and any remainder it has stays available to the agent's other requests.</li>
<li><strong>It meters against the completing token.</strong> The held invocation is metered against the auth token presented on the poll that completes it. The resource applies <xref target="overshoot"/> to that token before executing: it reserves the invocation's maximum cost against that token's remainder, serves, and commits.</li>
<li><strong>A completing token that does not fit does not complete it.</strong> If the invocation's maximum cost exceeds the remainder of the token presented on the poll, the resource does not execute it. It answers the poll with another <tt>202</tt> carrying <tt>requirement=auth-token</tt>, <tt>reason=insufficient-budget</tt>, and a fresh resource token whose <tt>presented_jti</tt> names that token, and continues to hold the invocation.</li>
<li><strong>A replay is not metered again.</strong> A resource answering a repeated presentation of the completing token from the stored result, as the base protocol requires, MUST NOT meter the invocation a second time. The <tt>AAuth-Budget</tt> on the replay reports that token's current <tt>remaining</tt>, so that the figure is still a floor, with the invocation's <tt>cost</tt> where the resource has it and no <tt>reserved</tt>.</li>
</ol>
<t><tt>AAuth-Budget</tt> on the <tt>202</tt> reports what it would report on the <tt>401</tt>: the <tt>remaining</tt> of the token presented on the request, <tt>required</tt> under <tt>insufficient-budget</tt> <xref target="required-member"/>, and <tt>unit</tt> and <tt>decimals</tt>. It carries no <tt>cost</tt> and no <tt>reserved</tt>, because nothing was served or held against that grant. A later <tt>202</tt> answering a poll that bore an auth token reports on that token the same way <xref target="header-rules"/>. The response to the completing poll carries <tt>AAuth-Budget</tt> for the completing token, as any response to a request bearing it does.</t>
<t>Under <tt>insufficient-budget</tt>, the agent's other move — lower the request's bound and retry on the token it already holds <xref target="exhaustion"/> — abandons the pending URL. The agent sends a new request with the smaller bound and does not poll the pending URL again. The held invocation is never executed, and since nothing was reserved for it (rule 1), abandoning it costs the grant nothing. The resource discards it when the pending URL's lifetime ends (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Pending URL Security).</t>
</section>

<section anchor="exhaustion-boundaries"><name>Boundaries</name>
<t><tt>429 Too Many Requests</tt> is not used by this document. The base protocol gives it two meanings already — <tt>slow_down</tt> on a pending URL, which asks a polling agent to lengthen its interval, and <tt>rate_limited</tt> at a revocation endpoint, which refuses work a caller has sent too much of (<xref target="I-D.hardt-oauth-aauth-protocol"/>). Neither is a budget condition. A budget refusal is about what the next request would cost, not about how often requests arrive, and the agent answers it by re-authorizing rather than by waiting.</t>
<t><tt>402 Payment Required</tt> is a different condition: the resource needs payment rather than re-authorization from the person. The base protocol already permits <tt>AAuth-Requirement</tt> on a <tt>402</tt>, and this document does not change that.</t>
</section>
</section>

<section anchor="enforcement"><name>Enforcement</name>

<section anchor="resource-enforces"><name>The Resource Enforces</name>
<t>The PS authorizes a number. The resource counts. The PS is not in the request path and does not meter.</t>
</section>

<section anchor="requires-auth-token"><name>Budgets Require Auth-Token Mode</name>
<t>A budget is carried in the <tt>budget</tt> claim of an auth token, so a resource can enforce one only where it requires auth tokens. A resource that declares <tt>access_mode: person-token</tt> and serves requests on the person's identity alone (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Person Identity Access) has no auth token to read a budget from, and neither does a resource operating in <tt>agent-token</tt> or <tt>session-token</tt> mode.</t>
<t>A metered resource therefore requires an auth token on the endpoints it meters (<tt>access_mode: auth-token</tt>). It MAY continue to serve unmetered endpoints on a person token, since a resource MAY apply different modes to different endpoints (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Resource Access Modes). A resource MUST NOT rely on this extension for an endpoint it serves without an auth token.</t>
<t>Where the resource holds the authorization state itself rather than reading it from a signed claim, there is nothing for the person server to have bounded.</t>
</section>

<section anchor="token-scope"><name>Token Scope</name>
<t>A budget is scoped to the auth token that carries it and expires with it. There is no persistent grant identifier and no requirement that the PS carry a budget across re-issuance. This is the mechanism, not a gap: re-issuance is where the PS re-decides <xref target="ps-decision"/>, and a budget that survived it would be a standing grant the PS no longer sizes.</t>
<t>The budget is revoked with the token. Revoking an auth token at its resource ends its budget along with the rest of its authorization. In three-party access the PS revokes it there; in four-party access the PS revokes the person token at the AS, and the AS revokes the auth tokens it issued against it (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Token Revocation). A resource without a revocation endpoint honors a revoked token until its <tt>exp</tt>. Consumption already committed is unaffected — a budget is a ceiling on spending, not a claim on what was spent — and a request already in flight completes, because revocation stops a token being used again rather than interrupting a call. This document adds nothing to that mechanism; it is named here because a person hitting stop expects the money to stop, and expiry alone bounds that at an hour. What a revocation means for the issuer's own accounting — when the allocation it reserved can be released — is <xref target="settlement"/>.</t>
<t>Two conditions return the agent to the PS, and either is sufficient. The auth token nears its <tt>exp</tt>: the base protocol caps an auth token at one hour and has the agent refresh it when fewer than five minutes remain (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Expiry and the Refresh Margin). Or its budget is exhausted <xref target="exhaustion"/>, which happens after however much work it took to spend. Expiry is proportional to time and exhaustion is proportional to spend, so the supervision interval tracks whichever is moving faster: a mission running cheaply reports on the hour, one running expensively reports in minutes, and no party configures the difference.</t>
</section>

<section anchor="unreported-allocations"><name>Unreported Allocations</name>
<t>An allocation can expire before its issuer sees a figure for it. A consumption record comes back only on a challenge to a request that presented the token <xref target="budget-consumed"/>. A token the agent retires by refreshing it, as the base protocol has it do inside the refresh margin <xref target="exhaustion"/>, is never named in one, and neither is a token held by an agent that has crashed or been abandoned. Until the next usage query, the issuer knows a grant was live and nothing about what it consumed — the true figure is anywhere from zero to the full <tt>amount</tt>.</t>
<t>Until a figure arrives, the issuer MUST account for the allocation as fully consumed. The error modes are not symmetric: an issuer that assumes less than was spent sizes the next allocation against headroom that may not exist and can overrun the ceiling it holds, while one that assumes the maximum is at worst temporarily conservative, which the next figure corrects. This is <xref target="ambiguous-failure"/> applied to the other end of the grant — the party missing a figure assumes the maximum until a real one supersedes it.</t>
<t>Reconciliation is idempotent where the channel names the token. A consumption record is a <tt>{jti, consumed}</tt> pair whose <tt>consumed</tt> is that token's total as of the resource token that carried it <xref target="budget-consumed"/>, so a record arriving late — or arriving again — replaces the assumption for that token rather than adding to it. Usage counters and per-key figures <xref target="usage-counters"/> name no <tt>jti</tt>; they are aggregates the issuer reconciles against its own ledger of what it assumed.</t>
<t>The rule falls on every party that sized the allocation against a ceiling it holds. In three-party access that is the PS <xref target="ps-decision"/>. In four-party access it is also the AS, which issues against the ceiling the PS stated <xref target="as-token-endpoint"/> and may apply credit or risk limits of its own <xref target="narrowing-chain"/>. Both are entitled callers of the usage endpoint <xref target="usage-authorization"/>.</t>

<section anchor="settlement"><name>Settlement</name>
<t>The issuer set every allocation's <tt>exp</tt>, so it knows when the assumption can be settled.</t>
<t>A consumption record settles nothing. It is a snapshot <xref target="budget-consumed"/>: the token stays valid, the agent may spend more on it, and under <tt>insufficient-budget</tt> may retry a smaller request on it <xref target="exhaustion"/>. What the record gives the issuer is the spend so far, for the re-authorization decision it is making at that moment.</t>
<t>A usage reading settles every allocation, in aggregate. The issuer takes the person's figure from the usage endpoint <xref target="usage-counters"/>, complete through <tt>as_of</tt>, and then holds</t>

<artwork><![CDATA[free = ceiling − metered − Σ amount of every allocation not yet settled
]]>
</artwork>
<t>where <tt>metered</tt> is the person's counter — <tt>all_time</tt> for a standing ceiling, the matching calendar counter for a calendar one — and an allocation is settled by the reading once its <tt>exp</tt> is at or before <tt>as_of</tt>. A settled allocation needs no figure of its own: whatever it consumed is inside <tt>metered</tt>, and it is no longer reserved. The only over-count in <tt>free</tt> is consumption under still-live tokens, present in both terms; it vanishes as each expires and the next reading covers it. The per-token figure in a consumption record refines the issuer's picture of a live allocation between readings; it does not settle it.</t>
<t>A revoked token is not an early <tt>exp</tt>. Revocation stops the spending; it does not by itself release the reservation, because the figure for what was spent still arrives with the next reading. What it does is move the moment after which no more can be spent, and the issuer learns whether that moment exists from the revocation's own answer: a recipient answers <tt>200</tt> once its cascade is terminal, directly or at the pending URL of a <tt>202</tt>, and an AS reports each resource's outcome in <tt>downstream</tt> (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Revocation Response). Three outcomes, three answers:</t>

<ul spacing="compact">
<li><strong>Recorded at the resource.</strong> Nothing further can be spent on that token. The allocation is settled by the first usage reading whose <tt>as_of</tt> is at or after the revocation was recorded, rather than the first one at or after <tt>exp</tt> — the same rule as any other settlement, since whatever was spent is inside <tt>metered</tt> by then.</li>
<li><strong><tt>revocation_unsupported</tt>.</strong> The resource honors the token until <tt>exp</tt>. The allocation stays reserved until <tt>exp</tt> and settles as it would have without the revocation.</li>
<li><strong><tt>revocation_unavailable</tt>.</strong> The issuer cannot tell which of the two it has, so it holds the allocation reserved until <tt>exp</tt> and MAY revoke again later.</li>
</ul>
<t>An issuer that treated every revocation as an immediate settlement would release headroom that a resource with no revocation endpoint is still spending against, which is <xref target="unreported-allocations"/> in reverse and has the error mode that section rules out.</t>
<t>An issuer whose ceiling is per calendar period SHOULD set each allocation's <tt>exp</tt> no later than the period boundary. No allocation then straddles two periods, every allocation of a period has expired when the period ends, and the period settles without a query. The cost is that a token issued near the boundary is short. The calendar counters have no sub-day period <xref target="calendar-counters"/>; an issuer with an hourly ceiling and clipped allocations never needs one, and an issuer with a trailing window settles from differences of successive <tt>all_time</tt> readings.</t>
</section>
</section>

<section anchor="aggregation"><name>Aggregation</name>
<t>Two things are counted, against different keys, and they are not the same requirement.</t>
<t><strong>The cap the resource enforces is per auth token.</strong> It is the <tt>budget</tt> claim of the token presented, and <xref target="overshoot"/> states the invariant: committed consumption plus outstanding reservations against <em>that token</em> MUST NOT exceed <em>its</em> granted <tt>amount</tt>. A resource needs no cross-token arithmetic to enforce a budget.</t>
<t><strong>The ledger the resource keeps is per person.</strong> The resource MUST aggregate consumption against the key <tt>(ps, sub, aud)</tt> of the auth token, which is what the consumption record <xref target="budget-consumed"/> and the usage counters <xref target="usage-counters"/> report. <tt>(ps, sub)</tt> identifies the person: <tt>sub</tt> is minted by the person server named in <tt>ps</tt>, and values from different person servers are different people. <tt>aud</tt> is the resource itself. In three-party access <tt>ps</tt> equals <tt>iss</tt>. In four-party access <tt>iss</tt> is the AS, which did not mint <tt>sub</tt>, so a key on <tt>iss</tt> would let two person servers behind one AS collide; <tt>ps</tt> is the person server the AS verified sent the token request (<xref target="I-D.hardt-oauth-aauth-protocol"/>, PS-to-AS Token Request). This document introduces no new identifier.</t>
<t>The ledger is not a second ceiling. A resource MUST NOT refuse a request that fits its token's budget because a per-person total has reached some figure the resource inferred; no party told it such a figure, and the budgets it was handed are what it was authorized to honor. Holding a person's spending across concurrent tokens within bounds is the PS's job <xref target="concurrency"/>, because the PS is the party that issues them and the only one that knows the ceiling <xref target="ps-decision"/>.</t>
<t>A per-agent ceiling is not a resource-side key either. A person server that wants one agent capped at less than another issues it a smaller allocation <xref target="ps-decision"/>; the enforcement is the token's own budget, and no resource-side dimension is involved. What a resource cannot supply from allocations alone is how much each agent actually spent, since an allocation is a ceiling rather than a figure — that is what the per-key query at the usage endpoint serves <xref target="per-key"/>.</t>
<t>The ledger's key is the person, not the agent and not the mission. An auth token names no agent, and a person's spending at a resource is theirs whichever agent incurred it; a per-agent key would also reset every time the person changed agents. <tt>mission_s256</tt> is optional — a token may carry one or not — so a mission-keyed ledger has no bucket for a mission-less token, and <xref target="inference"/> requires mission-less tokens for standing inference budgets. The person is the only key present on every auth token. Missions are an attribution dimension over that ledger <xref target="mission-attribution"/>, not the ledger itself.</t>
</section>

<section anchor="billing-account"><name>The Billing Account</name>
<t>A resource that meters usually charges someone for it, and the party it charges is an account in its own systems. Nothing in a budget names that account. The aggregation key above is <tt>(ps, sub, aud)</tt>, and <tt>sub</tt> is directed per person server — it identifies a person at one PS and carries no meaning at the resource beyond what the resource has learned about it.</t>
<t>For most resources that is sufficient and no mechanism is needed. The base protocol keys a person's relationship with a resource on the person server and <tt>sub</tt> precisely so it survives a change of agent, and a resource holding one account per person looks the account up from <tt>(ps, sub)</tt>. <tt>tenant</tt> names the person's organization and is not part of the identifier (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Organization Identification). Consumption then meters against the account the resource already had.</t>
<t>Beyond that, two different questions arise, and they compose rather than substitute. The first is asked once per person; the second on every authorization.</t>
<t><strong>Which person is this?</strong> The first budgeted request carrying a <tt>sub</tt> the resource has not seen is a question for the person, not the agent, and both access modes answer it with an interaction the person completes at the party that holds the account.</t>
<t>In three-party access the resource asks. It puts an <tt>interaction</tt> claim in the resource token, and the person server chains the person through the resource's own flow — signing in, creating an account, connecting a payment method — before completing its own consent (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Resource-Initiated Interaction). Because the resource issues the resource token, it decides when to ask again: once per person, or once per mission, since it sees <tt>mission_s256</tt> at that moment.</t>
<t>In four-party access the AS asks, returning <tt>202</tt> with <tt>requirement=interaction</tt> to the person server's token request (<xref target="I-D.hardt-oauth-aauth-protocol"/>, PS-AS Trust Establishment). This is the same one-time binding the AS already performs to establish trust with a person server, answering a second question at the moment it is already asking the person who they are.</t>
<t>Because <tt>sub</tt> is directed per person server, a person reaching the same resource through two person servers presents two identifiers. The binding interaction is what attaches both to one account, and a resource that skips it sees two people and bills two ledgers.</t>
<t><strong>Which of their accounts?</strong> Binding establishes who the person is. It does not say which of several accounts an authorization is for, and a person who holds more than one at the resource has to say. Account Binding (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Account Binding) carries the answer: an OPTIONAL <tt>account</tt> parameter on the authorization endpoint request, named from the resource's own namespace, echoed as the <tt>account</tt> claim of the resource token and copied into the auth token.</t>
<t>This applies in both access modes; how <tt>account</tt> reaches the issuer, and what each party does with it, is specified there and not restated here. In four-party access it reaches the AS in the resource token, so the binding tells the AS who the person is and <tt>account</tt> tells it which of their accounts this authorization bills.</t>
<t>A metered resource should ask for <tt>account</tt> where a person may hold more than one, because a budget enforced against the wrong account is charged to the wrong payer. A resource holding one balance per person needs none of it: binding is the whole mechanism, and <tt>account</tt> never appears in its tokens.</t>
<t>None of this is specific to budgets, and this document defines no new mechanism for it. It is stated here because a metered resource meets these cases on its first request and the rest of this document is silent on them.</t>
</section>

<section anchor="concurrency"><name>Concurrency</name>
<t>An agent may hold several concurrent auth tokens at the same resource — the <tt>mission_s256</tt> claim means concurrent missions produce concurrent tokens, each with its own budget, for up to an hour. Handling this is mandatory, not optional:</t>

<ul spacing="compact">
<li>A resource MUST apply the reserve-commit-release invariant of <xref target="overshoot"/> atomically per auth token, so that concurrent requests presenting the same token cannot together exceed its budget.</li>
<li>A resource MUST post consumption to the <tt>(ps, sub, aud)</tt> ledger <xref target="aggregation"/> atomically, so that concurrent requests across different tokens do not lose or double-count against the record and counters.</li>
<li>A PS SHOULD size per-token budgets so that their sum stays within whatever standing ceiling it holds for the person at that resource. This is the only place the cross-token total is enforced; <xref target="settlement"/> is how an issuer keeps that sum exact as allocations expire.</li>
</ul>
<t>The bound on over-issuance is the auth token lifetime multiplied by the number of concurrent tokens. A PS that issues <em>n</em> concurrent tokens of <em>X</em> each has authorized up to <em>nX</em> for as long as an hour, regardless of any standing figure it intended to hold.</t>
</section>

<section anchor="overshoot"><name>The Budget Is a Hard Cap</name>
<t>A resource MUST NOT let metered consumption exceed the granted <tt>amount</tt>. A budget is an authorization, and an authorization the enforcer may exceed is a hint.</t>
<t>Output is metered after it is generated, so honoring the cap means bounding the request before serving it. Before serving a chargeable request, the resource determines the request's maximum cost — from a bound the request declares, such as a maximum output length, or from a documented default — and refuses the request with <tt>reason=insufficient-budget</tt> <xref target="reason-parameter"/> when that maximum exceeds the remainder. An operation with no finite cost bound MUST be given one, be truncated when the remainder is consumed, or be refused.</t>
<t>The implementation shape is reserve-commit-release: atomically reserve the maximum against the grant, serve, commit the actual charge, release the difference. The normative requirement is the invariant, not the mechanism: committed consumption plus outstanding reservations MUST NOT exceed the granted <tt>amount</tt> at any moment, including under concurrent requests <xref target="concurrency"/>.</t>
<t>The honest cost of the invariant lands near exhaustion: a request bounded at more than the remainder is refused even when its actual cost would have fit. The mitigation is the agent's — read <tt>remaining</tt> from the refusal's <tt>AAuth-Budget</tt> header and retry with a bound that fits — and that is the correct pressure, since it rewards realistic bounds.</t>
</section>

<section anchor="failed-calls"><name>Failed Calls</name>
<t>Whether a request that fails — a <tt>5xx</tt> after input tokens were consumed — draws down the budget is the resource's metering policy. Whatever it meters, the record and counters MUST reflect, or the person's numbers will not reconcile with the bill.</t>
</section>

<section anchor="delegation"><name>Delegation</name>
<t>A budget is not delegable. Each auth token carries its own budget, and consumption through a delegated token draws down that token's budget alone. Both delegation paths — call chaining and sub-agent authorization (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Agent Delegation) — obtain each downstream auth token through the person server, so the PS sizes every grant in a chain, and <xref target="concurrency"/> already governs their sum. No sub-budget or draw-down-from-parent mechanism is defined.</t>
</section>

<section anchor="mission-attribution"><name>Mission Attribution</name>
<t>Consumption is attributed to a mission using the <tt>mission_s256</tt> claim of the auth token, beside its <tt>ps</tt> claim: a mission is identified by the PS that approved it together with <tt>mission_s256</tt> (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Mission Identifier). <tt>mission_s256</tt> is PS-asserted throughout: the person server validated the mission when it issued the person token, the resource copied it into the resource token, and the person server, and in four-party access the AS, verified that it matches the presented token's (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Resource Token Verification). No party in the chain takes the mission on the agent's word, which is what makes a resource's per-mission figures worth reading.</t>
</section>
</section>

<section anchor="usage-counters"><name>Usage Counters</name>
<t>The consumption record reaches the issuer only when the agent brings a resource token back. The agent is the party being budgeted and also the courier of the evidence: it cannot falsify the record, but between re-authorizations it does not appear, and the issuer is blind for up to an hour per token. The <tt>usage_endpoint</tt> <xref target="budget-units"/> removes the agent from that loop: the issuer — the PS, or in four-party access the AS — queries the resource directly, on a channel the agent is never on.</t>
<t>The endpoint serves <strong>usage counters</strong>: pre-summed consumption totals the PS reads, acts on, and displays. It does not compute, convert, or round. The resource keeps a handful of running integers, incremented at metering time; serving the endpoint requires no per-record history.</t>

<section anchor="usage-request"><name>Usage Request</name>
<t>The caller — a person server, or in four-party access an access server — MUST make a signed POST to the <tt>usage_endpoint</tt>, signing as a server in its own right: an HTTP Sig under the <tt>jwks_uri</tt> scheme, with <tt>id</tt> its <tt>issuer</tt> and <tt>dwk</tt> its metadata document (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Keying Material), and the signature additionally covering <tt>content-type</tt> and <tt>content-digest</tt>, as on a request with a body to a PS or AS (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Covered Components).</t>
<t>The body carries at most one <strong>scope key</strong>, naming a claim value the resource has seen in auth tokens:</t>

<ul spacing="compact">
<li><strong><tt>sub</tt></strong>: A directed person identifier. Scope: the person at this resource, across all their agents and missions.</li>
<li><strong><tt>tenant</tt></strong>: A tenant identifier. Scope: the organization, across its people.</li>
<li><strong><tt>mission_s256</tt></strong>: A mission identifier. Scope: one mission.</li>
</ul>
<t>and one OPTIONAL member:</t>

<ul spacing="compact">
<li><strong><tt>jkts</tt></strong>: An array of JWK Thumbprints (<xref target="RFC7638"/>), each naming a signing key the resource has seen present an auth token. Asks for what each of those keys consumed <xref target="per-key"/>.</li>
</ul>
<t>An access server names the person server whose values it is asking about:</t>

<ul spacing="compact">
<li><strong><tt>ps</tt></strong>: The person server whose namespace the scope key belongs to. REQUIRED when the caller is an access server and the request carries a scope key, since the tokens an AS issues carry values from every person server that federates with it. A person server omits it; it is the caller.</li>
</ul>
<t>A request MUST carry a scope key or <tt>jkts</tt>, and MAY carry both. At most one scope key may appear. A request with more than one scope key, or with neither a scope key nor <tt>jkts</tt>, is an error <xref target="usage-authorization"/>.</t>

<sourcecode type="http"><![CDATA[POST /usage HTTP/1.1
Host: inference.example
Content-Type: application/json
Content-Digest: sha-256=:...:
Signature-Input: sig=("@method" "@authority" "@path"
    "content-type" "content-digest"
    "signature-key");created=1754620000
Signature: sig=:...signature bytes...:
Signature-Key: sig=jwks_uri;id="https://ps.example";
    dwk="aauth-person.json";kid="key-1"

{
  "sub": "8f14e45fceea167a5a36dedd4bea2543",
  "jkts": ["NzbLsXh8uDCcd-6MNwXF4W_7noWXFZAfHkxZsRGC9Xs",
           "0ZcOCORZNYy-DWpqq30BbmLzO1Yw3ZQhIgnHZQKxNVE"]
}
]]>
</sourcecode>
</section>

<section anchor="usage-response"><name>Usage Response</name>

<sourcecode type="json"><![CDATA[{
  "as_of": 1754619970,
  "aud": "https://ps.example",
  "unit": "USD",
  "decimals": 6,
  "sub": "8f14e45fceea167a5a36dedd4bea2543",
  "usage": {
    "day": 1243180,
    "week": 3118400,
    "month": 8432650,
    "year": 39847220,
    "all_time": 61438050
  },
  "jkts": {
    "NzbLsXh8uDCcd-6MNwXF4W_7noWXFZAfHkxZsRGC9Xs": 38215600,
    "0ZcOCORZNYy-DWpqq30BbmLzO1Yw3ZQhIgnHZQKxNVE": 23222450
  }
}
]]>
</sourcecode>

<ul spacing="compact">
<li><strong><tt>as_of</tt></strong> (REQUIRED): The time through which the figures are complete, in seconds since the Unix epoch. Metering aggregation MAY lag serving; <tt>as_of</tt> is what keeps a lagging figure honest.</li>
<li><strong><tt>aud</tt></strong> (REQUIRED): The caller the response was produced for — a person server, identified as in the <tt>ps</tt> claim of a resource token, or an access server, identified as in the <tt>iss</tt> claim of an auth token. It is what stops a signed response being presented to a third party as a statement about them <xref target="signed-response"/>.</li>
<li><strong><tt>unit</tt></strong> (REQUIRED) and <strong><tt>decimals</tt></strong> (REQUIRED): The unit every figure in the response is denominated in, and its scale, as in the budget object <xref target="budget-object"/>.</li>
<li>The <strong>scope key</strong> from the request, echoed unchanged — <tt>sub</tt>, <tt>tenant</tt>, or <tt>mission_s256</tt> — present only when the request carried one.</li>
<li><strong><tt>usage</tt></strong> (REQUIRED when the request carried a scope key): Calendar counters for that scope.</li>
<li><strong><tt>jkts</tt></strong> (REQUIRED when the request carried <tt>jkts</tt>): An object mapping each thumbprint to what that key consumed <xref target="per-key"/>.</li>
</ul>

<section anchor="calendar-counters"><name>Calendar Counters</name>
<t>The members of <tt>usage</tt> are <tt>day</tt>, <tt>week</tt>, <tt>month</tt>, <tt>year</tt>, and <tt>all_time</tt>, following the interval enumeration of Stripe Issuing <xref target="Stripe.Issuing"/> <xref target="pa-stripe"/>. Each is a non-negative integer giving consumption within the current period:</t>

<ul spacing="compact">
<li><strong><tt>day</tt></strong>: since 00:00 UTC today.</li>
<li><strong><tt>week</tt></strong>: since Monday 00:00 UTC of the current ISO 8601 week.</li>
<li><strong><tt>month</tt></strong>, <strong><tt>year</tt></strong>: since the start of the current UTC calendar month or year.</li>
<li><strong><tt>all_time</tt></strong> (REQUIRED): everything the resource has metered against budgets under this key.</li>
</ul>
<t>All period boundaries are UTC. This is a definition, not a deployment choice: no timezone appears in metadata or in the response, and every party computes the same figure. The cost is that "today" resets mid-afternoon in Auckland, which Stripe Issuing accepts for the same reason this document does — the counter is decision context, not a bill.</t>
<t>The calendar counters other than <tt>all_time</tt> are OPTIONAL; a resource omits periods it does not track. For a <tt>mission_s256</tt> query, <tt>all_time</tt> is the mission total — the figure a PS wants when deciding whether to fund a mission's continuation — and a resource MAY serve it alone.</t>
<t>Counters are subject to the 15-digit bound of <xref target="range-limits"/>. A resource whose cumulative figure would exceed it omits that counter rather than reporting an inexact number.</t>
<t><tt>all_time</tt> reaches as far back as the resource retains. This document sets no retention requirement for scope keys; the person's bill is the durable record.</t>
</section>

<section anchor="per-key"><name>Per-Key Figures</name>
<t>Each member of <tt>jkts</tt> is a thumbprint mapped to a single non-negative integer: everything the resource has metered against budgets on auth tokens presented by that key.</t>
<t>There are no calendar periods here. A key's consumption is already bounded by the tokens issued to it, and an auth token lives at most an hour; a key that has stopped presenting tokens has a figure that no longer moves. Periods answer "how much this month", which is a question about a person, not about a key.</t>
<t>A resource SHOULD retain a key's figure for at least 24 hours after that key's last metered request, and MAY retain it longer. The bound is idle time rather than age, so a key in continuous use is never pruned. Beyond that window the PS is the party that accumulates: it polls, it knows which keys belonged to which agent across rotations, and it holds the history. The resource keeps a short tail.</t>
<t>A resource MUST omit a thumbprint from <tt>jkts</tt> rather than report zero for it when it holds no figure — because the key is unrecognized, or because its figure has been pruned. Absence means the resource cannot answer; a present zero means the key consumed nothing. This differs from the treatment of an unrecognized scope key <xref target="usage-authorization"/>, and the reason is that there is nothing to conceal: the PS issued or relayed every auth token, so it already knows the key exists, and a zero that means "pruned" would be a wrong answer to an allocation decision rather than a withheld one.</t>
</section>

<section anchor="one-unit"><name>One Unit Per Response</name>
<t>Every figure in a response is in one unit, named once at the top level, and it is the unit the resource meters in. There is no request parameter selecting it.</t>
<t>The alternative is a per-unit array at every level, which costs every response the shape needed by deployments that meter in one unit — which is nearly all of them, since a resource that meters several quantities collapses them to one billing unit before denominating a budget <xref target="non-goals"/>. A resource may still declare several units in <tt>budget_units</tt> <xref target="budget-units"/>, because that is what an agent may ask a budget to be denominated in; what this endpoint reports is what the resource actually metered, and the response says which unit that was.</t>
</section>

<section anchor="signed-response"><name>The Signed Response</name>
<t>Signing the usage response is RECOMMENDED. A resource that signs uses an HTTP Sig with a key from the <tt>jwks_uri</tt> in its resource metadata <xref target="budget-units"/> — the same key material the PS already fetched to verify resource tokens. The signature MUST cover <tt>@status</tt>, <tt>content-type</tt>, and <tt>content-digest</tt>, and MUST be bound to the request by covering the request's <tt>@authority</tt> and <tt>@path</tt> with the <tt>req</tt> parameter (<xref target="RFC9421"/>).</t>

<sourcecode type="http"><![CDATA[HTTP/1.1 200 OK
Content-Type: application/json
Content-Digest: sha-256=:...:
Signature-Input: sig=("@status" "content-type"
    "content-digest" "@authority";req "@path";req);
    created=1754620001
Signature: sig=:...signature bytes...:
Signature-Key: sig=jwks_uri;id="https://inference.example";
    dwk="aauth-resource.json";kid="r1"
]]>
</sourcecode>
<t>The endpoint reports what a person owes for, so an unsigned figure is one the party that produced it can later disown. Signing makes the resource committed to what it reported: it cannot tell the person server one number and the biller another. It does not make the meter honest — the resource is the counterparty as well as the signer — and <xref target="counters-trust"/> covers what remains.</t>
<t>It also closes an asymmetry. The consumption record <xref target="budget-consumed"/> is already resource-signed, because it rides inside a resource token. The usage endpoint is the only issuer-facing consumption channel that is not.</t>
<t>It is RECOMMENDED rather than REQUIRED because the figures are decision context rather than authorization, and because this would be the first response-side signature in the AAuth family — the signature profile is request-side throughout <xref target="header-trust"/>. A person server receiving an unsigned response is not in a position to do anything but read it: refusing it leaves the PS with no figures rather than unattributable ones, which is the worse of the two. What signing changes is whether the resource can later disown what it said, and that is worth having wherever both ends will implement it.</t>
<t><tt>aud</tt> is what keeps the signed response non-transferable. Without it a resource-signed statement of consumption could be handed to a third party as though it described them, and the figures carry no other indication of who asked.</t>
</section>
</section>

<section anchor="usage-authorization"><name>Authorization and Errors</name>
<t>The <tt>id</tt> parameter of the <tt>Signature-Key</tt> header is the caller's server identifier, and its <tt>dwk</tt> names the metadata document that says which role is calling — <tt>aauth-person.json</tt> for a person server, <tt>aauth-access.json</tt> for an access server (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Keying Material). That identifier is what the response carries as <tt>aud</tt> <xref target="usage-response"/>. A caller is a person server or an access server. The resource MUST only answer for values that have appeared in auth tokens it accepted whose <tt>iss</tt> or <tt>ps</tt> claim names the caller: for a PS, the tokens it issued in three-party access and the tokens carrying it as <tt>ps</tt> in four-party access; for an AS, the tokens it issued, and for a scope key only those whose <tt>ps</tt> is the request's <tt>ps</tt>. This applies to thumbprints in <tt>jkts</tt> as much as to scope keys. <tt>sub</tt> is directed per PS, so one person server cannot even name another's subjects; <tt>tenant</tt>, <tt>mission_s256</tt>, and thumbprints are not directed, and this check is what stops a third party from querying them.</t>
<t>An AS is entitled because it sizes allocations against a ceiling of its own <xref target="narrowing-chain"/> and is bound by <xref target="unreported-allocations"/> for them. An AS operated by the resource may take the same figures from the resource directly; the endpoint is for the AS that is not.</t>
<t>A query for a scope key the resource does not recognize returns <tt>200</tt> with <tt>usage</tt> omitted; "never seen" and "nothing consumed" are deliberately indistinguishable, so that a query cannot be used to discover whether a person holds an account. Unrecognized thumbprints are handled differently and for a stated reason <xref target="per-key"/>.</t>
<t><tt>invalid_request</tt>, using the error response format of (<xref target="I-D.hardt-oauth-aauth-protocol"/>), is returned for a body carrying more than one scope key, carrying neither a scope key nor <tt>jkts</tt>, carrying a scope key from an access server without <tt>ps</tt>, or carrying a malformed value.</t>
<t>A resource MAY rate-limit the endpoint, using the <tt>RateLimit</tt> fields (<xref target="I-D.ietf-httpapi-ratelimit-headers"/>) as on any endpoint. A PS SHOULD poll no faster than its decisions require.</t>
</section>

<section anchor="usage-division"><name>Division of Labor</name>
<t>The two issuer-facing channels answer different questions at different moments. The consumption record <xref target="budget-consumed"/> serves the re-authorization decision: it arrives in-band, resource-signed, at no round-trip cost, exactly when the issuer is deciding. The usage counters serve everything else: supervision between re-authorizations, settlement of allocations that expired unreported <xref target="settlement"/>, mission totals, tenant-level exposure, how a person's spending divides among their agents <xref target="per-key"/>, and the person's dashboard. A metered resource SHOULD implement both; the narrowing chain <xref target="narrowing-chain"/> functions with the record alone. The agent's own view is neither of these: it is the <tt>AAuth-Budget</tt> header <xref target="aauth-budget-header"/>, scoped to the token it holds and to the request it just made.</t>
</section>
</section>

<section anchor="capability"><name>Capability Negotiation</name>
<t>This document adds <tt>budget</tt> to the AAuth Capability Value Registry. An agent that understands budget semantics — the <tt>budget</tt> claim, the <tt>AAuth-Budget</tt> header, and the <tt>reason</tt> values <xref target="reason-parameter"/> — SHOULD include <tt>budget</tt> in its <tt>AAuth-Capabilities</tt> request header, and in the <tt>capabilities</tt> parameter of its person token and auth token requests (<xref target="I-D.hardt-oauth-aauth-protocol"/>, AAuth-Capabilities Request Header). The person token request is where a budget-aware agent first reaches its PS about a resource, so it is where the declaration first matters:</t>

<sourcecode type="http"><![CDATA[AAuth-Capabilities: interaction, clarification, budget
]]>
</sourcecode>
<t>This tells the resource up front whether the agent will read what it sends, rather than leaving it to rely on ignore-unknown-parameter behavior. It is the same negotiation <tt>interaction</tt> performs for <tt>requirement=interaction</tt>. Recipients MUST ignore unrecognized capability values.</t>
<t>A resource MUST NOT withhold <tt>AAuth-Budget</tt> from an agent that did not declare the capability. The header is safe to ignore.</t>
</section>

<section anchor="errors"><name>Errors</name>
<t>Narrowing is never an error. A resource that will grant less than the agent asked for issues a smaller <tt>budget</tt>; a PS or AS that will allow less issues a smaller <tt>amount</tt>. No error is returned in either case.</t>
<t>Errors are reserved for statements that cannot be reconciled:</t>
<table>
<thead>
<tr>
<th>Error</th>
<th>Status</th>
<th>Endpoint</th>
<th>Meaning</th>
</tr>
</thead>

<tbody>
<tr>
<td><tt>invalid_budget</tt></td>
<td>400</td>
<td>Authorization endpoint</td>
<td>The <tt>budget</tt> object is malformed, or names a <tt>unit</tt> the resource has not declared in <tt>budget_units</tt></td>
</tr>

<tr>
<td><tt>invalid_budget</tt></td>
<td>400</td>
<td>PS and AS auth token endpoints</td>
<td>The resource token's <tt>budget.decimals</tt> disagrees with the value declared for that unit in the resource's <tt>budget_units</tt> metadata, or the object is otherwise malformed</td>
</tr>

<tr>
<td><tt>invalid_request</tt></td>
<td>400</td>
<td>Usage endpoint</td>
<td>More than one scope key, neither a scope key nor <tt>jkts</tt>, or a malformed value <xref target="usage-authorization"/></td>
</tr>
</tbody>
</table><t>Error responses use the error response format defined in AAuth Protocol (<xref target="I-D.hardt-oauth-aauth-protocol"/>).</t>
<t><tt>invalid_budget</tt> at the authorization endpoint parallels <tt>invalid_scope</tt>: the agent named something the resource does not recognize.</t>
<t>A <tt>decimals</tt> mismatch at the PS or AS MUST be a hard reject rather than a narrowing. A mismatch means one of the two figures is stale, and guessing which one is worse than failing — the error modes are a budget interpreted a thousandfold too large or too small.</t>
<t>A PS or AS that has no cached <tt>budget_units</tt> for the resource and cannot fetch the metadata MAY proceed on the resource token's <tt>decimals</tt>, which is signed by the resource. The cross-check is a defense against staleness, not against the resource.</t>

<section anchor="declining"><name>Declining with Guidance</name>
<t>A PS granting less than the resource offered needs no mechanism: the <tt>amount</tt> in the issued claim says it. A PS or AS that declines a token request outright on budget grounds MAY include <strong><tt>suggested_budget</tt></strong> in its error response body: a budget object <xref target="budget-object"/> whose <tt>unit</tt> and <tt>decimals</tt> are copied from the resource token, and whose <tt>amount</tt> is what the issuer would currently accept. It is guidance, not a grant — the agent's move is a fresh resource token at that figure <xref target="authorization-endpoint"/> and a new token request, with the usual consent and policy evaluation.</t>
<t>In four-party access an AS's <tt>suggested_budget</tt> reaches the PS, not the agent. The PS relays it in the problem+json body with which it relays the AS's error to the agent (<xref target="I-D.hardt-oauth-aauth-protocol"/>, Auth Token Delivery), with <tt>unit</tt> and <tt>decimals</tt> unchanged and <tt>amount</tt> lowered to what the PS would itself currently accept for the person where that is lower. That figure never exceeds the ceiling the PS holds <xref target="ps-decision"/>, and the PS states it as an amount it would accept, not as the ceiling, which the agent does not learn <xref target="why-not-the-ceiling"/>.</t>
</section>
</section>

<section anchor="inference"><name>Standing Authorization for Metered Inference</name>
<t>Metered inference is the initiating use case for this extension and has a property that distinguishes it from most resource access: the agent needs it before it can do anything else, including decide what else it needs.</t>
<t><strong>Inference authorization is agent-scoped and standing.</strong> It is established at first PS contact, not per mission. The reason is a termination argument rather than a preference: a per-mission inference budget requires the agent to run inference in order to evaluate whether it needs more inference budget, and that evaluation itself consumes inference. A standing allocation terminates; a per-mission one does not.</t>
<t><strong>Mission-less auth tokens are already legal.</strong> <tt>mission_s256</tt> is optional in a person token and therefore in everything derived from it, and the PS permission and interaction endpoints work with or without a mission. No new token type is needed for this pattern.</t>
<t><strong>Do not model the harness as a mission.</strong> A mission <tt>description</tt> is defined as human intent expressed in Markdown. Synthesizing an "operating mission" or "harness mission" to hold the inference budget would put a fabricated description into the content-addressed blob that every token carrying <tt>mission_s256</tt> is bound to, and mission revocation — a kill switch for one piece of work — would then also be the kill switch for the agent's ability to think. Those are two distinct controls and conflating them is a mistake.</t>
<t><strong>Pre-mission inference runs on a mission-less auth token</strong> and lands in the person's usage with no mission attribution. Mission boundaries become natural auth token re-issuance points, which costs nothing: the agent is already talking to its PS to create the mission.</t>
<t><strong>AP-bundled inference is out of scope.</strong> Where the agent provider pays for the agent's inference, the agent token is the credential, the person server never appears, and nothing in this document applies. This is the common deployment today, and readers will otherwise assume the extension covers it. The case that does engage this extension is a harness written by the agent provider spending against the <em>person's</em> inference account — which is where a ceiling matters most, because the agent provider's software is deciding how much of the person's money to consume.</t>
<t><strong>A TPX provider adds agent support without touching its meter.</strong> A TPX <xref target="TPX"/> provider already prices per token, reports cost in each response's <tt>usage</tt>, and enforces a person-granted budget as a hard cap — for human-driven apps holding OAuth grants. Serving agents means accepting AAuth auth tokens carrying <tt>budget</tt> beside those grants: two authorization envelopes over one metering core. Deployed this way, this extension is the AAuth binding of TPX.</t>
<t>Neither envelope displaces the other. A person driving an app and an agent acting for that person are different situations, and a provider serving both has one meter under two front doors <xref target="implementation-status"/>.</t>
</section>

<section anchor="security-considerations"><name>Security Considerations</name>

<section anchor="header-trust"><name>The Header Is Unsigned</name>
<t><tt>AAuth-Budget</tt> is not signed, so an intermediary can lie about the balance. The failure modes are bounded. Understating <tt>remaining</tt> makes the agent re-authorize earlier than it needed to. Overstating it makes the agent hit an unexpected <tt>401</tt>. Neither causes overspend, because enforcement is the resource checking metered consumption against the signed <tt>budget</tt> claim of the auth token — the header is a pacing signal, not the authorization.</t>
<t>An implementation that needs the balance to be trustworthy rather than merely harmless can cover <tt>AAuth-Budget</tt> with an HTTP Message Signature on the response (<xref target="RFC9421"/>), as the usage endpoint recommends for its own responses <xref target="signed-response"/>. It is not required here, and the difference is what each response is for. A usage response is a statement of what a person owes for, read by a party that may later have to hold the resource to it. A balance on a request response is a pacing signal the agent acts on immediately and re-reads on the next call, and signing every metered response to protect a figure that is superseded seconds later buys little for what it costs at inference volumes.</t>
</section>

<section anchor="budget-not-scope"><name>Budget Is Not a Substitute for Scope</name>
<t>A budget bounds how much an agent may consume, not what it may do. An agent with a small budget at a resource that performs irreversible actions can still perform them. Resources MUST continue to enforce scope, and SHOULD NOT treat the presence of a budget as evidence that the person reviewed anything beyond the amount.</t>
</section>

<section anchor="security-concurrency"><name>Over-Issuance Through Concurrency</name>
<t>A PS that issues concurrent auth tokens without tracking their sum has authorized their sum, not any single one of them <xref target="concurrency"/>. A resource that aggregates non-atomically across concurrent requests can be driven past the ceiling by parallel calls. Both are implementation errors that produce real overspend, and both are easy to make.</t>
</section>

<section anchor="counters-trust"><name>Consumption Reports as Attack Surface</name>
<t>The <tt>budget_consumed</tt> record is resource-signed and usage counters are served from the resource's authenticated endpoint, signed where the resource follows <xref target="signed-response"/>; both are the resource's own account of what it metered. A resource that inflates them can induce a PS to authorize more than the person intended, or to refuse further authorization. An issuer SHOULD reconcile them against the person's billing relationship with the resource where one exists, and SHOULD NOT treat them as authoritative for anything other than its own next decision.</t>
<t>Signing the usage response <xref target="signed-response"/> does not change this. It makes the resource committed to a figure rather than able to disown it, which is what stops it reporting one number to the person server and another to the biller. It does not make the meter honest, because the resource meters, reports, and signs. The person's bill is the record a dispute settles against, and a signed report is evidence of what the resource said, not of what it consumed.</t>
</section>

<section anchor="unit-substitution"><name>Unit Substitution</name>
<t>Because a resource may change the unit between what the agent requested and what it offers <xref target="narrowing-chain"/>, a PS reading a <tt>budget</tt> claim must read <tt>unit</tt> and <tt>decimals</tt> rather than assuming the denomination the agent described in <tt>justification</tt>. The cross-check against <tt>budget_units</tt> in <xref target="errors"/> is the defense against a stale scale; there is no defense against a resource that misdenominates deliberately, and none is needed — that resource is equally free to ignore the budget it was granted.</t>
</section>
</section>

<section anchor="privacy-considerations"><name>Privacy Considerations</name>
<t>A budget amount is financial information about the person. It travels from the resource to the PS in the resource token and back in the auth token, and appears in plaintext in the <tt>AAuth-Budget</tt> response header on every response.</t>
<t>The consumption record and usage counters are more revealing than the budget itself: counters describe the person's spending at that resource over time and by mission, and the record itemizes one grant. All of it is visible to the person's PS by design — the PS is deciding on the person's behalf.</t>
<t>The two channels differ in who else can read them. Usage counters travel only between resource and issuer, so tenant-scope and long-horizon figures exist nowhere the agent can see. The consumption record rides in the resource token, which the agent relays and can read. It names only the token the agent itself presented, and only as a <tt>jti</tt> and an amount — an agent learns nothing about the person's other tokens or other agents from it — and a resource that considers even that too revealing omits <tt>budget_consumed</tt> and serves the usage endpoint alone.</t>
<t>Because <tt>AAuth-Budget</tt> is unsigned and unencrypted above TLS, every intermediary on the path sees the person's remaining balance at that resource, and the price of the request that produced the response. Carrying no cumulative figure bounds this: an intermediary sees what one call cost and what is left of one hour's grant, not the person's spending history at that resource. Deployments that consider even the per-call figure sensitive may omit the OPTIONAL <tt>cost</tt> member from the header, at the price of leaving the agent to derive it from successive <tt>remaining</tt> values.</t>
</section>

<section anchor="iana-considerations"><name>IANA Considerations</name>

<section anchor="http-header-field-registration"><name>HTTP Header Field Registration</name>
<t>This specification registers the following HTTP header field in the "Hypertext Transfer Protocol (HTTP) Field Name Registry" established by <xref target="RFC9110"/>:</t>

<ul spacing="compact">
<li>Header Field Name: <tt>AAuth-Budget</tt></li>
<li>Status: permanent</li>
<li>Structured Type: Dictionary</li>
<li>Reference: This document, <xref target="aauth-budget-header"/></li>
</ul>
</section>

<section anchor="jwt-claims-registration"><name>JWT Claims Registration</name>
<t>This document requests registration of the following claims in the IANA "JSON Web Token Claims" registry established by <xref target="RFC7519"/>:</t>
<table>
<thead>
<tr>
<th>Claim Name</th>
<th>Claim Description</th>
<th>Change Controller</th>
<th>Reference</th>
</tr>
</thead>

<tbody>
<tr>
<td><tt>budget</tt></td>
<td>Authorized spending ceiling at a resource</td>
<td>IETF</td>
<td>This document, <xref target="budget-object"/></td>
</tr>

<tr>
<td><tt>budget_consumed</tt></td>
<td>What the presented auth token consumed, reported by a resource</td>
<td>IETF</td>
<td>This document, <xref target="budget-consumed"/></td>
</tr>
</tbody>
</table></section>

<section anchor="aauth-capability-value-registry"><name>AAuth Capability Value Registry</name>
<t>This document requests registration of the following value in the AAuth Capability Value Registry established by AAuth Protocol (<xref target="I-D.hardt-oauth-aauth-protocol"/>). The registry policy is Specification Required (<xref target="RFC8126"/>, Section 4.6).</t>
<table>
<thead>
<tr>
<th>Value</th>
<th>Reference</th>
</tr>
</thead>

<tbody>
<tr>
<td><tt>budget</tt></td>
<td>This document, <xref target="capability"/></td>
</tr>
</tbody>
</table></section>

<section anchor="no-budget-unit-registry"><name>No Budget Unit Registry</name>
<t>This document deliberately establishes no registry of unit values. Units are declared by each resource in its <tt>budget_units</tt> metadata, exactly as scope values are declared in <tt>scope_descriptions</tt>. See <xref target="why-no-unit-registry"/>.</t>
</section>
</section>

<section anchor="implementation-status"><name>Implementation Status</name>
<t><em>Note: This section is to be removed before publishing as an RFC.</em></t>
<t>This section records the status of known implementations of the protocol defined by this specification at the time of posting of this Internet-Draft, and is based on a proposal described in <xref target="RFC7942"/>. The description of implementations in this section is intended to assist the IETF in its decision processes in progressing drafts to RFCs.</t>

<section anchor="implementations-in-progress"><name>Implementations in Progress</name>
<t><strong>tokenpony</strong> (Infinite Logic PBC) meters LLM inference and implements TPX <xref target="TPX"/>, an OAuth 2.0 profile carrying the same grant for human-driven apps. That deployment is live, and this extension is being added over the same metering core, with AuthGravity as the person server and Harness News as the agent, in four-party access. TPX is a complete deployment on its own: it needs no person server and nothing from this document, and the two specifications share no wire surface — one meter, two independent authorization envelopes.</t>
<t><strong>Regent Protocol</strong> is in production at get4agent.com (marketplace resource and provisioning server): allocations, refusals with <tt>required</tt>, the streaming cost-omitted mode, the usage endpoint, and sub-agent delegation. Its open-source resource/PS middleware <tt>regent-httpsig</tt> (Python, Apache-2.0) publishes test vectors.</t>
<t><strong>The editor</strong> is implementing this extension in several services.</t>
<t>Implementation reports and test vectors are expected from these efforts and will be recorded here.</t>
</section>
</section>

<section anchor="document-history"><name>Document History</name>
<t><em>Note: This section is to be removed before publishing as an RFC.</em></t>

<ul spacing="compact">
<li><t>draft-hardt-aauth-budgets-00</t>

<ul spacing="compact">
<li>Initial submission, written against draft-hardt-oauth-aauth-protocol-11.</li>
</ul></li>
</ul>
</section>

<section anchor="acknowledgments"><name>Acknowledgments</name>
<t>The author thanks Abay Aubakirov, Alex Polvi, and Karl McGuinness for feedback on early drafts.</t>
</section>

</middle>

<back>
<references><name>References</name>
<references><name>Normative References</name>
<reference anchor="I-D.hardt-httpbis-signature-key" target="https://datatracker.ietf.org/doc/draft-hardt-httpbis-signature-key">
  <front>
    <title>HTTP Signature Keys</title>
    <author fullname="Dick Hardt" initials="D." surname="Hardt">
      <organization>Hellō</organization>
    </author>
    <author fullname="Thibault Meunier" initials="T." surname="Meunier">
      <organization>Cloudflare</organization>
    </author>
    <date year="2026"/>
  </front>
</reference>
<reference anchor="I-D.hardt-oauth-aauth-protocol" target="https://datatracker.ietf.org/doc/draft-hardt-oauth-aauth-protocol">
  <front>
    <title>AAuth Protocol</title>
    <author fullname="Dick Hardt" initials="D." surname="Hardt">
      <organization>Hellō</organization>
    </author>
    <date year="2026"/>
  </front>
</reference>
<xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.7519.xml"/>
<xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.7638.xml"/>
<xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.8126.xml"/>
<xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.9110.xml"/>
<reference anchor="RFC9111" target="https://www.rfc-editor.org/info/rfc9111">
  <front>
    <title>HTTP Caching</title>
    <author fullname="R. Fielding" initials="R." surname="Fielding" role="editor"/>
    <author fullname="M. Nottingham" initials="M." surname="Nottingham" role="editor"/>
    <author fullname="J. Reschke" initials="J." surname="Reschke" role="editor"/>
    <date year="2022" month="June"/>
  </front>
  <seriesInfo name="STD" value="98"/>
  <seriesInfo name="RFC" value="9111"/>
  <seriesInfo name="DOI" value="10.17487/RFC9111"/>
</reference>
<xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.9421.xml"/>
<reference anchor="RFC9651" target="https://www.rfc-editor.org/info/rfc9651">
  <front>
    <title>Structured Field Values for HTTP</title>
    <author fullname="M. Nottingham" initials="M." surname="Nottingham"/>
    <author fullname="P-H. Kamp" surname="P-H. Kamp"/>
    <date year="2024" month="September"/>
  </front>
  <seriesInfo name="RFC" value="9651"/>
  <seriesInfo name="DOI" value="10.17487/RFC9651"/>
</reference>
</references>
<references><name>Informative References</name>
<reference anchor="AP2" target="https://ap2-protocol.org/">
  <front>
    <title>Agent Payments Protocol (AP2)</title>
    <author>
      <organization>Google</organization>
    </author>
    <date year="2025"/>
  </front>
</reference>
<reference anchor="ERC20" target="https://eips.ethereum.org/EIPS/eip-20">
  <front>
    <title>EIP-20: Token Standard</title>
    <author fullname="Fabian Vogelsteller" initials="F." surname="Vogelsteller"/>
    <author fullname="Vitalik Buterin" initials="V." surname="Buterin"/>
    <date year="2015"/>
  </front>
</reference>
<reference anchor="I-D.hardt-aauth-r3" target="https://datatracker.ietf.org/doc/draft-hardt-aauth-r3">
  <front>
    <title>AAuth Rich Resource Requests (R3)</title>
    <author fullname="Dick Hardt" initials="D." surname="Hardt">
      <organization>Hellō</organization>
    </author>
    <date year="2026"/>
  </front>
</reference>
<reference anchor="I-D.ietf-httpapi-ratelimit-headers" target="https://datatracker.ietf.org/doc/html/draft-ietf-httpapi-ratelimit-headers-11">
  <front>
    <title>RateLimit header fields for HTTP</title>
    <author fullname="Roberto Polli" initials="R." surname="Polli">
      <organization>Team Digitale, Italian Government</organization>
    </author>
    <author fullname="Alex Martínez Ruiz" initials="A. M." surname="Ruiz">
      <organization>Red Hat</organization>
    </author>
    <author fullname="Darrel Miller" initials="D." surname="Miller">
      <organization>Microsoft</organization>
    </author>
    <date year="2026" month="May" day="23"/>
  </front>
  <seriesInfo name="Internet-Draft" value="draft-ietf-httpapi-ratelimit-headers-11"/>
</reference>
<reference anchor="ISO4217" target="https://www.iso.org/iso-4217-currency-codes.html">
  <front>
    <title>ISO 4217:2015 Codes for the representation of currencies</title>
    <author>
      <organization>International Organization for Standardization</organization>
    </author>
    <date year="2015"/>
  </front>
</reference>
<reference anchor="OBIE.VRP" target="https://openbankinguk.github.io/read-write-api-site3/v4.0/profiles/vrp-profile.html">
  <front>
    <title>Variable Recurring Payments Profile, Read/Write Data API</title>
    <author>
      <organization>Open Banking Implementation Entity</organization>
    </author>
    <date year="2024"/>
  </front>
</reference>
<reference anchor="ODRL" target="https://www.w3.org/TR/odrl-vocab/">
  <front>
    <title>ODRL Vocabulary and Expression 2.2</title>
    <author>
      <organization>W3C</organization>
    </author>
    <date year="2018"/>
  </front>
</reference>
<xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.7942.xml"/>
<xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.9396.xml"/>
<xi:include href="https://bib.ietf.org/public/rfc/bibxml/reference.RFC.9457.xml"/>
<reference anchor="Stripe.Issuing" target="https://docs.stripe.com/api/cards/object#card_object-spending_controls">
  <front>
    <title>Stripe Issuing: Card spending controls</title>
    <author>
      <organization>Stripe</organization>
    </author>
    <date year="2026"/>
  </front>
</reference>
<reference anchor="TPX" target="https://tokenpony.dev/spec/">
  <front>
    <title>TPX: Token Pony Express — An OAuth 2.0 Profile for Metered LLM Inference Grants</title>
    <author>
      <organization>Infinite Logic PBC</organization>
    </author>
    <date year="2026"/>
  </front>
</reference>
<reference anchor="W3C.PaymentRequest" target="https://www.w3.org/TR/payment-request/">
  <front>
    <title>Payment Request API</title>
    <author>
      <organization>W3C</organization>
    </author>
    <date year="2025"/>
  </front>
</reference>
<reference anchor="x402" target="https://docs.x402.org">
  <front>
    <title>x402: HTTP 402 Payment Protocol</title>
    <author>
      <organization>x402 Foundation</organization>
    </author>
    <date year="2025"/>
  </front>
</reference>
</references>
</references>

<section anchor="design-rationale"><name>Design Rationale</name>

<section anchor="why-integer"><name>Why an Integer in a Declared Scale</name>
<t><strong>Why not a Structured Field Decimal.</strong> <xref target="RFC9651"/>, Section 3.3.2 caps a Decimal at 12 integer digits and 3 fractional digits, with round-half-to-even serialization. Inference is priced per million tokens, so a single call can cost on the order of $0.000015, which serializes to <tt>0.000</tt>. The header would report zero consumption on every request while the person receives a bill.</t>
<t><strong>Why not a binary float.</strong> IEEE 754 cannot represent 0.1 exactly, and a budget is a running sum, so the error accumulates rather than cancelling. No monetary prior art surveyed in <xref target="prior-art"/> uses a float.</t>
<t><strong>Why not a decimal string,</strong> which is what most monetary prior art does use (<xref target="W3C.PaymentRequest"/>, <xref target="x402"/>). An integer in a declared scale is the only representation under which the Structured Field header value and the JWT claim value are the same number. A decimal string in the claim beside a Structured Field Decimal in the header would disagree at the third decimal place, and every implementation would have to define its own rounding to reconcile them.</t>
</section>

<section anchor="why-explicit-scale"><name>Why the Scale Is Explicit</name>
<t>x402 <xref target="x402"/> omits a scale because the asset contract address <em>is</em> the unit and the scale comes from the contract, immutably. <tt>USD</tt> does not work that way. ISO 4217 <xref target="ISO4217"/> fixes the minor unit for USD at 2, and 2 decimal places cannot express a fifteen-microdollar call. The resource is therefore choosing a scale that the unit identifier does not determine, and a scale that is chosen has to travel with the amount.</t>
</section>

<section anchor="why-signed-decimals"><name>Why <tt>decimals</tt> Is in the Signed Claim</name>
<t><tt>decimals</tt> is declared in <tt>budget_units</tt> metadata and also carried in the <tt>budget</tt> claim. The duplication is deliberate, and the reason is not primarily an attack argument — the resource is the enforcer, and a resource willing to reinterpret the scale is equally willing to ignore the budget outright.</t>
<t>The reasons are versioning and self-description:</t>

<ul spacing="compact">
<li><strong>Versioning.</strong> Auth tokens live up to an hour. A resource changing its declared precision as an ordinary product decision — cents to micro-dollars — would silently reinterpret every token in flight by a factor of ten thousand.</li>
<li><strong>Self-description.</strong> The resource is not the only reader. The PS ledgers against the number, the person's dashboard displays it, a proxy logs it, a dispute cites it. Each would otherwise have to resolve the number against mutable metadata fetched at some unspecified later time.</li>
</ul>
</section>

<section anchor="why-one-record"><name>Why One Record</name>
<t>An earlier revision carried up to twenty consumption records in the resource token, one per recent grant, so that a PS re-deciding saw the person's recent spend at the resource without a round trip. Working through the issuer's ledger, the presented token's figure is the only one that is news. Every prior token either came back through the same path when it was retired, in which case its record already arrived, or expired unreported, in which case the conservative rule <xref target="unreported-allocations"/> holds until a usage reading settles it <xref target="settlement"/> — and a reading settles every expired allocation at once, without naming any of them. The prior records cost roughly a kilobyte in the <tt>401</tt> header and let an agent read what was spent under tokens it never held, including the person's other agents'. The one thing the list uniquely offered — sibling-token spend at a checkpoint — is what the per-key query at the usage endpoint answers, per agent rather than per token, on a channel the agent is not on.</t>
<t>The <tt>jti</tt> stays. Concurrent tokens are the normal case for an agent running several missions, and an issuer holding several live allocations for one agent cannot tell from an amount alone which of them a figure belongs to. The alternative — attributing by agent key and falling back to the conservative rule when ambiguous — makes the fallback the normal case for exactly the agents doing concurrent work. One short claim keeps the figure attributable and idempotent per token.</t>
</section>

<section anchor="why-no-cumulative"><name>Why the Agent Is Not Told Its Cumulative Consumption</name>
<t>An earlier revision of this document carried a <tt>consumed</tt> member in the header, giving what had been consumed against the auth token to date, and a <tt>balance_endpoint</tt> where the agent could read the same figure on demand. Both are gone. The agent is told what a request cost and what is left; cumulative consumption is reported to the person server and not to the agent.</t>
<t>Four reasons.</t>
<t>The figure is redundant. <tt>granted</tt> is in the auth token the agent signed the request with, and <tt>remaining</tt> is in the response, so cumulative consumption is <tt>granted - remaining</tt> whenever nothing is reserved. Carrying a member the recipient can already compute is weight without information.</t>
<t>The word is already spoken for. A consumption record is a <tt>{jti, consumed}</tt> pair <xref target="budget-consumed"/> where <tt>consumed</tt> means the total metered against one token's whole budget. A person server reads those records and this header's semantics against each other. Having <tt>consumed</tt> mean a per-request figure in one place and a per-token total in the other is the kind of collision that survives review and then costs an implementer a day.</t>
<t>Cumulative spend is the more revealing figure. It describes a pattern rather than a transaction, and <tt>AAuth-Budget</tt> travels unsigned past every intermediary on the path <xref target="privacy-considerations"/>. What the agent genuinely needs per response is the price of the call it just made and whether it can afford another. Both are per-request facts, and that is what the field now carries.</t>
<t>There is a fourth reason that applies across tokens rather than within one. An agent holding its cumulative consumption over a series of allocations can watch the series and infer the ceiling behind it — how much the PS is willing to release, and how fast. The ceiling is deliberately not disclosed <xref target="why-not-the-ceiling"/>, and a per-token figure that reconstructs it by subtraction discloses it anyway.</t>
<t>The <tt>balance_endpoint</tt> went with it. It was OPTIONAL, existed for one case — an ambiguous failure, where the agent cannot say whether a request was metered — and cost a resource an endpoint to implement and this document a section to specify. That case is now answered by a rule the agent applies locally <xref target="ambiguous-failure"/>: assume the maximum until a later <tt>remaining</tt> says otherwise. A conservative default that every agent applies is better than an optional round trip that some resources offer.</t>
</section>

<section anchor="why-not-the-ceiling"><name>Why the Granted Budget Is Not the Person's Ceiling</name>
<t>A person server could authorize the whole of a person's intended spend at a resource in one auth token and let the agent draw it down. TPX <xref target="TPX"/> does the OAuth equivalent: the budget sits on a durable grant, the app spends against it unsupervised, and the person hears about it when the grant runs dry. This document does not, and the difference is not a missing feature.</t>
<t><strong>The figure is not knowable when it would have to be fixed.</strong> A mission is approved before the work is done, and the work is what determines the cost. A person asked at approval for a number is guessing. Too low and the agent stops mid-task and the person is interrupted anyway. Too high and the number is not a control, because the agent will never reach it and nothing is checked before it does. An allocation sized against what the work has actually cost so far does not require the guess to be right.</t>
<t><strong>A ceiling the agent can read is a ceiling the agent plans against.</strong> An agent that knows it has been authorized for a pool treats the pool as available. An agent that knows only its current allocation asks when the allocation runs out, and asking is what puts the PS back in the decision. The <tt>justification</tt> accompanying that request, and the consumption records arriving with it, are the person server's evidence — and neither exists if the agent never has to come back.</t>
<t><strong>The check-in is the point, not a cost of it.</strong> Re-authorization is where the PS reads what the last allocation bought <xref target="ps-inputs"/>, compares it against the mission, and chooses among its six responses <xref target="ps-responses"/> — including the two that are not a number at all: asking the agent to account for the spend, and ending the work. A single up-front grant has no such moment. It has one, at approval, when the least is known.</t>
<t>This is why the ceiling appears nowhere on the wire. There is no claim for it, it may not be shared with the agent, and cumulative consumption that would reveal it by subtraction is withheld as well <xref target="why-no-cumulative"/>. What the resource enforces is the allocation in the token in front of it. What the person authorized is a matter between the person and their PS.</t>
</section>

<section anchor="why-trailer-adds"><name>Why a Trailer Only Adds a Member</name>
<t><tt>AAuth-Budget</tt> may be sent as a trailer, carrying <tt>cost</tt> for a response whose cost was unknown when the header was written <xref target="streaming"/>. This document rejected trailers in an earlier revision, on the reasoning of the RateLimit work (<xref target="I-D.ietf-httpapi-ratelimit-headers"/>): intermediaries drop them, and combining a header value with a trailer value complicates clients.</t>
<t>The first objection carries less weight here. It is largely a browser property — <tt>fetch</tt> does not surface trailers — and the traffic this document governs is an agent calling a resource, which is server-to-server and commonly HTTP/2. Where a trailer is dropped anyway, rule 3 of <xref target="trailer-rules"/> makes that harmless.</t>
<t>The second objection is the real one, and <xref target="trailer-rules"/> answers it structurally rather than by mandating client behavior. A recipient may discard trailer fields or merge them into the header set, and a sender can neither observe nor control which. Restating a member across the two would therefore produce a value that depends on the recipient's choice: <tt>AAuth-Budget</tt> is a Dictionary, duplicate keys resolve last-wins (<xref target="RFC9651"/>, Section 3.2), so merging yields the trailer's value and discarding yields the header's. Adding a member that the header did not carry has no such fork. Merge yields a complete picture, discard yields a conservative one, and neither is wrong.</t>
<t>This is why <tt>reserved</tt> is a statement about a request rather than a running balance. A running figure would go stale the moment the reservation was released, and a stale member surviving a merge is exactly the failure the rule exists to prevent. What was held for a request does not change after the fact.</t>
</section>

<section anchor="why-not-ratelimit"><name>Why Not <tt>RateLimit</tt></name>
<t><tt>RateLimit</tt> and <tt>RateLimit-Policy</tt> (<xref target="I-D.ietf-httpapi-ratelimit-headers"/>) already report a server-side quota and its remaining balance. Reusing them would avoid a new header. Four things prevent it, all following from a budget being an authorization rather than a capacity hint:</t>

<ol spacing="compact">
<li><strong>Stated non-goal.</strong> The RateLimit specification excludes authorization from its scope. Reporting the balance of a PS-issued grant through a field whose own specification says it is not for access control is a misuse a reviewer will name.</li>
<li><strong>No unit carrier.</strong> <tt>q</tt> and <tt>r</tt> MUST be non-negative Integers, and the quota units registry covers <tt>request</tt>, <tt>content-bytes</tt>, and <tt>concurrent-requests</tt>. There is no currency carrier, so the denomination would be invisible in the field reporting the number.</li>
<li><strong>Opposite reliability contracts.</strong> RateLimit says servers need not send the fields on every response, clients must not assume future responses will carry them, and a positive <tt>r</tt> is not a guarantee of anything. Those are correct properties for a capacity hint and wrong ones for the remaining portion of an authorization, which is why <xref target="header-rules"/> requires the field on every response the metering layer answers rather than leaving it optional.</li>
<li><strong>Intermediary rewriting.</strong> Intermediaries MAY tighten <tt>RateLimit</tt> values. An intermediary tightening a budget balance is forging authorization state.</li>
</ol>
<t>The two fields are complementary and MAY appear on the same response. A resource limiting an agent to 100 requests per minute <em>and</em> to five dollars of spend is stating two different things, and collapsing them loses one.</t>
</section>

<section anchor="why-not-mission-aggregate"><name>Why Budget Is Not a Mission Aggregate</name>
<t>A mission description may well say "budget around $5,000," and the obvious next step is to make that a protocol object the PS enforces across every resource the mission touches. This document does not, for two reasons.</t>
<t>The resources in a mission meter in units that do not add up. An airline meters in USD, an inference endpoint in micro-dollars or tokens, a storage service in gigabyte-months. Aggregating requires conversion rates and a common denomination, neither of which the protocol has, and both of which change.</t>
<t>More fundamentally, the enforcement point is wrong. The party that can enforce a per-resource ceiling is the resource, because it meters. No party meters the mission. A mission-level total is a PS policy input — the PS decides how much of the person's $5,000 to authorize at each resource as it goes — which is exactly what the narrowing chain in <xref target="narrowing-chain"/> gives it. The mission description already carries the person's intent, and the PS already reads it.</t>
</section>

<section anchor="why-no-unit-registry"><name>Why Units Are Not Registered</name>
<t>RateLimit establishes an IANA registry of quota units because its units — <tt>request</tt>, <tt>content-bytes</tt>, <tt>concurrent-requests</tt> — are protocol-generic and every server means the same thing by them.</t>
<t>Budget units are not generic. A unit is meaningful only against a resource's own pricing, and the parties that need to interpret it are the resource that declared it and the PS that fetched the resource's metadata. This is the same situation as scope values, which the base protocol leaves to each resource's <tt>scope_descriptions</tt> rather than registering. Registering <tt>USD</tt> would add nothing that ISO 4217 does not already provide, and registering <tt>tokens</tt> would suggest an interoperable meaning that does not exist.</t>
<t>No registry does not mean no constraint. A monetary unit SHOULD be an ISO 4217 alphabetic code <xref target="budget-object"/>, which is what keeps a consent screen able to render "$5.00" rather than a resource-invented string the person has to interpret. What is left unregistered is the non-monetary case, where no external register exists to point at.</t>
</section>
</section>

<section anchor="prior-art"><name>Prior Art</name>
<t>This section is non-normative. It records where the field shapes and value encodings in this document come from, and what gap remains.</t>

<section anchor="pa-rar"><name>RFC 9396, Rich Authorization Requests</name>
<t><xref target="RFC9396"/> (OAuth WG, May 2023), Section 2.2 defines the complete set of common data fields for an authorization detail: <tt>type</tt>, <tt>locations</tt>, <tt>actions</tt>, <tt>datatypes</tt>, <tt>identifier</tt>, <tt>privileges</tt>. There is no amount or quantity field among them. <tt>instructedAmount</tt> appears only in the document's examples and belongs to the <tt>payment_initiation</tt> authorization details type, which comes from Berlin Group NextGenPSD2 and carries ISO 20022 <tt>ActiveCurrencyAndAmount</tt> semantics; it is not defined by RFC 9396.</t>
<t>Section 10 states that registration of authorization details types with the AS is outside the specification's scope, and Section 14 registers the request parameter, the claim, the metadata fields, and the <tt>invalid_authorization_details</tt> error — but establishes no registry of types.</t>
<t>Section 6.1 is the load-bearing citation: there is no standardized mechanism for comparing two arbitrary authorization detail requests, and an AS should not rely on simple object comparison. AAuth Budgets closes that gap for one narrow case. Because the unit is resource-declared and the value is an integer in a declared scale, "is this grant less than the prior one" is a numeric comparison rather than a structural one.</t>
<t>R3 (<xref target="I-D.hardt-aauth-r3"/>) covers why AAuth does not profile RAR generally.</t>
</section>

<section anchor="pa-tpx"><name>TPX</name>
<t>TPX <xref target="TPX"/> is an OAuth 2.0 profile for metered LLM inference grants: apps ship without provider keys, and the person grants each app a metered budget from a provider the person chooses and pays. The budget rides RFC 9396 <tt>authorization_details</tt> as <tt>{type: "llm-inference", budget, models}</tt>; the user or provider MAY grant less than requested and the client MUST read the granted figure from the token response — the requested/granted split of the narrowing chain <xref target="narrowing-chain"/>, in OAuth form.</t>
<t>Three of its mechanisms have direct counterparts here. Its <tt>GET /credits</tt> spend summary and <tt>budget_used</tt> introspection member report cumulative spend on a grant, which this document reports to the person server <xref target="usage-counters"/> rather than to the agent <xref target="why-no-cumulative"/>. Its separation of <tt>budget_exhausted</tt> from <tt>balance_exhausted</tt> is the same boundary <xref target="exhaustion-boundaries"/> draws between re-authorization and payment. And its budget is "a damage cap, not a payment" — the hard-cap property <xref target="overshoot"/> states, which TPX asserts as a property of the grant and this document additionally realizes with reservations, since an agent's operations are unbounded before generation.</t>
<t>TPX denominates in USD as decimal JSON numbers with at most six fractional digits, where this document uses an integer in a declared scale <xref target="why-integer"/>. TPX serves human-driven apps over OAuth; <xref target="inference"/> describes the deployment where one provider's meter serves both envelopes.</t>
</section>

<section anchor="pa-vrp"><name>UK Open Banking Variable Recurring Payments</name>
<t><xref target="OBIE.VRP"/> <tt>ControlParameters</tt> is the most complete standing-budget model in production: <tt>MaximumIndividualAmount</tt>, <tt>MaximumCumulativeAmount</tt>, <tt>MaximumCumulativeNumberOfPayments</tt>, and <tt>PeriodicLimits[]</tt> carrying <tt>PeriodType</tt> and <tt>PeriodAlignment</tt>. <tt>PeriodAlignment</tt> — <tt>Consent</tt> versus <tt>Calendar</tt> — names the choice the usage counters <xref target="usage-counters"/> make: calendar alignment, with the timezone question answered by fixing UTC.</t>
</section>

<section anchor="pa-stripe"><name>Stripe Issuing Spending Controls</name>
<t><xref target="Stripe.Issuing"/> <tt>spending_controls.spending_limits[]</tt> is <tt>{amount, interval, categories}</tt>, with <tt>amount</tt> an integer in the currency's smallest unit. <tt>interval</tt> spans <tt>per_authorization</tt> through <tt>all_time</tt>, which is the precedent for expressing a per-transaction cap and a periodic cap in a single enumeration rather than two fields; the usage counters <xref target="usage-counters"/> adopt the date-based portion of that enumeration directly. Stripe computes all date-based intervals from midnight UTC — the precedent for fixing UTC by definition — and documents spending aggregation as best-effort with up to 30 seconds of delay, the precedent for <tt>as_of</tt>.</t>
</section>

<section anchor="pa-w3c"><name>W3C Payment Request API</name>
<t><xref target="W3C.PaymentRequest"/> <tt>PaymentCurrencyAmount</tt> is <tt>{currency, value}</tt> with <tt>value</tt> a decimal string. It is reused by Google's Agent Payments Protocol <xref target="AP2"/> in <tt>CartMandate</tt>. AP2's <tt>IntentMandate</tt> carries a maximum price, an expiry, and a merchant allowlist as user-signed constraints, which is the closest prior art to a user-authorized agent ceiling.</t>
</section>

<section anchor="pa-x402"><name>x402</name>
<t><xref target="x402"/> carries <tt>maxAmountRequired</tt> as a string in the asset's atomic units, alongside <tt>asset</tt> (a contract address) and <tt>network</tt>; v2 renames the field to <tt>amount</tt> and moves <tt>network</tt> to CAIP-2 form. There is no currency field: the asset identifier is the unit, and the exponent comes from the contract. This is the precedent for using a unit identifier rather than a currency field, and for <tt>decimals</tt> as the name of the scale.</t>
</section>

<section anchor="pa-odrl"><name>ODRL</name>
<t><xref target="ODRL"/> <tt>payAmount</tt> with <tt>unit</tt> is the precedent for a field literally named <tt>unit</tt> holding a currency code.</t>
</section>

<section anchor="pa-ratelimit"><name>RateLimit Header Fields</name>
<t><xref target="I-D.ietf-httpapi-ratelimit-headers"/> (HTTPAPI WG, Standards Track, not yet an RFC) defines <tt>RateLimit-Policy</tt> with <tt>q</tt>, <tt>qu</tt>, <tt>w</tt>, and <tt>pk</tt>, and <tt>RateLimit</tt> with <tt>r</tt>, <tt>t</tt>, and <tt>pk</tt>, both as RFC 9651 Lists. It establishes an IANA RateLimit Quota Units registry (Specification Required) with initial entries <tt>request</tt>, <tt>content-bytes</tt>, and <tt>concurrent-requests</tt>, and three RFC 9457 problem types: <tt>quota-exceeded</tt> (429), <tt>temporary-reduced-capacity</tt> (503), and <tt>abnormal-usage-detected</tt> (429).</t>
<t>It is cited here for why <tt>AAuth-Budget</tt> exists separately <xref target="why-not-ratelimit"/>, for the trailer rules <xref target="trailer-rules"/>, and for <tt>pk</tt> as the precedent for a documented, client-predictable partition key.</t>
</section>
</section>

</back>

</rfc>
