acceptodds
Under review as a conference paper at ICLR 2027

GRID: Graph Representation of Intelligence Data for Security Text Knowledge Graph Construction

Abstract

Security knowledge graphs can serve as computable and traceable external memory for security agents, yet existing LLMs lack domain knowledge grounded in real security text, and end-to-end document-to-graph training is hard to supervise with cheap and stable rewards. We present Graph Representation of Intelligence Data (G.R.I.D.), an end-to-end framework for security text knowledge graph construction. GRID builds security-domain supervision from CTI articles without manual labels, aligning each article with a traceable graph through graph extraction and knowledge-graph-conditioned text revision. It then reformulates documentto-graph learning as a scripted task bank of four-option multi-select questions and triple-level regex targets, which is cheaper and more stable than per-step LLM-judge scoring of full graphs. On a unified benchmark of 249 CTI articles from five sources (GRID, CASIE, CTINexus, MalKG, and SecureNLP), the posttrained Qwen3-4B-Instruct-2507 Task-bank Reward model reaches 84.62% sourceaveraged precision, 64.91% recall, and 68.53% Avg F1: the best recall and a near-tied top F1, with fewer tokens than the similarly strong CTINexus pipeline. The secondary End2End Reward model, trained with LLM-as-judge rewards, reaches 76.91% precision, 53.85% recall, and 58.06% Avg F1. The offline task bank, reusable across post-training runs, outperforms the online End2End reward, Choice-only Reward, and End2End SFT without RL, and both article rewriting and article-complexity ordering matter. Finally, this supervision is inexpensive to generate and reusable, in two independent senses: with the open-weight GPTOSS-120B generating the training data, the Qwen3-4B student still gains 13.8 absolute F1 points over its untrained baseline; and the task bank from our main experiments, reused with no new data generation, trains a different backbone, Llama-3.1-8B-Instruct, lifting it by 24.4 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.