acceptodds
Under review as a conference paper at ICLR 2027

Automated Hyperparameter Configuration Selection for Distantly Supervised NER using Reinforcement Learning and Bayesian Optimization

Abstract

Distantly supervised Named Entity Recognition (NER) produces noisy, incomplete labels that make training highly sensitive to hyperparameter choices. We study automated hyperparameter configuration selection for this setting, comparing Bayesian Optimization (Gaussian Process Thompson Sampling, GP-TS) and reinforcement learning (Twin Delayed DDPG, TD3; Soft Actor-Critic, SAC) against resource-aware hyperparameter optimization baselines (Random Search, Tree-structured Parzen Estimator, TPE) and hand-designed learning-rate schedules (Step Decay, Cosine Annealing, ReduceLROnPlateau). Each method selects a configuration using short-budget 2-epoch trials, after which the selected configuration is retrained for 10 epochs, so that every method is compared under an identical final training budget. Using BERT-CRF and RoBERTa-CRF under the BOND distant supervision framework, we evaluate on three benchmarks (Webpage, Wikigold, CoNLL-2003) with three random seeds and report meanstandard deviation. We find that configuration selection improves over a static baseline on every dataset: on CoNLL-2003, Random Search reaches 0.719 (RoBERTa-CRF) and 0.737 (BERT-CRF) against static baselines of 0.674 and 0.681, with GP-TS, TPE, TD3, and SAC all within 0.02; on Wikigold, TD3 and Random Search reach 0.480 and 0.481 against 0.452; on Webpage the gains are smaller but positive for the strongest selectors (SAC 0.558, Random Search 0.557 vs. 0.524). The hand-designed schedules do not help at this budget: Step Decay and ReduceLROnPlateau fall below the static baseline on Wikigold and CoNLL-2003, because neither reaches its first decay step within 10 epochs. We further show that search-stage scores from 2-epoch trials are unreliable proxies for 10-epoch retraining performance, and that best-checkpoint scores — the usual reporting convention — are indistinguishable across schedules here because the best checkpoint almost always falls in the first epoch or two, which is why we report final-epoch scores. Our results show that automated configuration selection is a viable, complementary strategy under weak supervision, and they provide practical guidance for evaluating such methods fairly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.