AI and especially large language models, face critical deployment challenges as inference con- sumes substantially more energy than training over their operational lifecycle. This thesis investigates efficient inference methodologies through three complementary approaches. First, we explore neural network scaling laws and training-inference compute trade-offs, exploring how Chinchilla scaling laws can be adapted to account for inference costs. Second, we analyze com- putational patterns in transformer architectures to understand optimization targets. Third, we conduct empirical studies comparing quantized and full-precision models across multiple scales, demonstrating that certain quantization methods and techniques maintain performance while reducing computational requirements. We extend those studies by examining test-time scaling behaviors of model, seeing if smaller models can more usefull than larger models if given extra compute at inference time. Our comprehensive analysis of weight and activation quantization establishes quantization as a fundamental technique for efficient deployment. This research provides both theoretical frameworks and practical methodologies for optimizing in- ference operations, enabling sophisticated language model deployment in resource-constrained environments while maintaining high performance.
AI and especially large language models, face critical deployment challenges as inference con- sumes substantially more energy than training over their operational lifecycle. This thesis investigates efficient inference methodologies through three complementary approaches. First, we explore neural network scaling laws and training-inference compute trade-offs, exploring how Chinchilla scaling laws can be adapted to account for inference costs. Second, we analyze com- putational patterns in transformer architectures to understand optimization targets. Third, we conduct empirical studies comparing quantized and full-precision models across multiple scales, demonstrating that certain quantization methods and techniques maintain performance while reducing computational requirements. We extend those studies by examining test-time scaling behaviors of model, seeing if smaller models can more usefull than larger models if given extra compute at inference time. Our comprehensive analysis of weight and activation quantization establishes quantization as a fundamental technique for efficient deployment. This research provides both theoretical frameworks and practical methodologies for optimizing in- ference operations, enabling sophisticated language model deployment in resource-constrained environments while maintaining high performance.
Training-Inference Resource Trade-offs for Sustainable AI Deployment / Šubić, T.. - (2026 Sep 21).
Training-Inference Resource Trade-offs for Sustainable AI Deployment
ŠUBIĆ, TOMISLAV
2026-09-21
Abstract
AI and especially large language models, face critical deployment challenges as inference con- sumes substantially more energy than training over their operational lifecycle. This thesis investigates efficient inference methodologies through three complementary approaches. First, we explore neural network scaling laws and training-inference compute trade-offs, exploring how Chinchilla scaling laws can be adapted to account for inference costs. Second, we analyze com- putational patterns in transformer architectures to understand optimization targets. Third, we conduct empirical studies comparing quantized and full-precision models across multiple scales, demonstrating that certain quantization methods and techniques maintain performance while reducing computational requirements. We extend those studies by examining test-time scaling behaviors of model, seeing if smaller models can more usefull than larger models if given extra compute at inference time. Our comprehensive analysis of weight and activation quantization establishes quantization as a fundamental technique for efficient deployment. This research provides both theoretical frameworks and practical methodologies for optimizing in- ference operations, enabling sophisticated language model deployment in resource-constrained environments while maintaining high performance.| File | Dimensione | Formato | |
|---|---|---|---|
|
SUBIC PhD Thesis.pdf
accesso aperto
Descrizione: SUBIC PhD thesis
Tipologia:
Tesi di dottorato
Dimensione
4.48 MB
Formato
Adobe PDF
|
4.48 MB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


