Foundation-model inference as a period cost
Why inference is usually expensed
Production inference is the cost of running the service, not building it. Under IAS 38 the day-to-day servicing of an asset and ongoing operating costs are not part of its cost and are expensed as incurred. The same logic sits behind the US GAAP treatment of post-implementation costs for internal-use software ASC 350-40. Metered inference serving live users is the archetype of a period cost.
When inference spend could be capitalised
The exception is inference consumed to build the asset itself: generating a fine-tuning dataset, running evaluation suites during development, or producing embeddings that become part of a delivered asset. These are development-phase activities, and their tokens may be capitalised if the recognition conditions are met and the phase is evidenced IAS 38 §54-62. The distinction is not the model or the endpoint; it is the purpose and the phase.
Build-phase versus run-phase tokens
- Build-phase: fine-tuning jobs, evaluation runs, dataset generation for a specific asset under development.
- Run-phase: serving production requests, ordinary retrieval and generation for live users.
- The phase marker on each tagged request, not the provider bill, decides the treatment.
The practical consequence
For most organisations the great majority of token spend is run-phase and therefore expensed. The capitalisable slice is real but narrow, and treating it as narrow is what keeps the asset defensible. Overreaching, by sweeping production inference into the balance sheet, is the fastest way to lose an audit challenge on the whole amount.