Retriever Optimization for Low-Resource Telecom Domain-Specific Question Answering
Abstract
Telecom question answering is a challenging setting for retrieval-augmented generation (RAG): evidence is fragmented across standards, research papers, encyclopedic resources, and domain web documents, while answers often depend on technical tables, equations, and specialized protocol language. In low-resource telecom subdomains, fully adapting the generator can over-specialize the model and degrade general capability, making retriever-centric adaptation especially attractive. To this end, we investigate which components should be optimized and which retriever training objectives are most effective for developing a domain-specific QA agent in a low-resource setting. We argue that retrieval is the most effective adaptation target compared with model fine-tuning. Theoretically, we show that when domain-specific data are scarce, retriever-side query-encoder tuning can generalize better than LoRA due to its lower estimation complexity. We then identify two particularly relevant retriever objectives: the latent-document RAG likelihood, which optimizes generation utility, and the contrastive InfoNCE objective, which improves semantic retrieval geometry. Rather than treating these objectives as static alternatives, we leverage them jointly through a retriever optimization method designed to maximize downstream QA performance in the telecom domain. Specifically, we introduce an adaptive temperature framework that inserts learnable temperatures into the retrieval distribution of the RAG objective and the softmax normalization of the InfoNCE objective, thereby shaping the gradient signal that each objective sends to the query encoder during training. We further add a query-distillation regularizer that constrains the learned query encoder to remain compatible with the frozen base document space, preserving the strong retrieval geometry of the base model while still allowing domain adaptation. Across telecom-specific retrieval and generative QA benchmarks, we show empirically that this proposed retriever optimization method improves both evidence retrieval and answer generation for a domain-specific RAG agent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.