CADMY Blog

Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion

by Fengzhuo Zhang, Zhuoran Yang, Dirk Bergemann  ·  September 9, 2026

Figure 1. Teaser map of the project. LLM personalization across many users forms a closed loop: users choose whether lightweight In-Context Learning (ICL) or compute-intensive Supervised Fine-Tuning (SFT) is attractive, based on the pretrained model, their desired task, and prevailing GPU congestion. These individual choices aggregate into shared GPU congestion, which in turn feeds back into the cost of personalization. Our contribution is to model this full loop: statistical personalization quality, equilibrium compute demand, and algorithm menus.

TL;DR

Personalization makes LLMs more useful, but it also consumes scarce shared compute. Once many users personalize models on the same infrastructure, their choices create congestion: one user’s fine-tuning job can increase the waiting cost faced by everyone else.

The main contribution of our paper is a tractable statistical–economic framework for this problem. It connects how ICL and SFT learn, how users choose between them under shared compute, and how platforms should price and design personalization menus. The main message is:

  • Algorithm performance is governed by the coverage and signal quality. ICL adapts within the topics covered by pretraining, while SFT can learn from user-specific data across both covered and uncovered topics. The linear model captures this distinction by separating the subspace covered by pretraining from the uncovered subspace. This separation reveals the key tradeoff between ICL and SFT: SFT dominates when the uncovered subspace carries sufficiently strong signal-to-noise; ICL dominates when those directions are too noisy to learn reliably.
  • Congestion turns the SFT–ICL comparison into an economic problem. Users each choose a personalization method and a sample size for personalization. These choices create aggregate compute demand for platforms, and we prove that the resulting equilibrium congestion level is uniquely pinned down, even when users can mix between algorithms.
  • Serving load can move non-monotonically with problem parameters. Through users’ equilibrium choices, the higher prior precision reduces the congestion level, but broader pretraining coverage, harder tasks, or more expensive SFT can raise or lower congestion through the interaction between sample intensity and algorithm switching.
  • Offering both ICL and SFT is profit-safe for the platform. From the platform’s perspective, raising congestion prices reduces equilibrium demand, and offering both ICL and SFT never lowers maximal platform profit in the model relative to a lightweight ICL-only menu. This helps explain why major LLM platforms increasingly offer both options.

Read the full post here.
 

Personalization Congestion

Figure 1. Teaser map of the project