May 2026 to May 2028

Inference cost and usage model

A twistable two-stream model for an AI-heavy developer: foreground work saturates at attention, background work rebounds into cheap tokens.

Today
-
tokens/day
Month 24
-
tokens/day
Today spend
-
USD/day, API only
Month 24 spend
-
with Jevons response

Tokens per day

foreground plus background

Foreground Background

API dollars per day

retail price band, no subscriptions

Month 24 sensitivity

one parameter moved +/-20%

tokens/day = H_fg q_fg k_fg tau_fg (1 + r) + J_bg d_bg tau_bg (1 + r)