acceptodds
Under review as a conference paper at ICLR 2027

SciSkillBank: A Dataset of Claude Code Skills for Scientific Research

Abstract

AI agents such as Claude Code follow reusable “skills,” instruction files (SKILL.md) that many authors now write for scientific research. No collection yet shows which research stages these skills delegate, whether they encode methodological knowledge, or how much control they leave to people. We release SciSkillBank, 2,200 research-relevant GitHub repositories (85.0% applied data science) that an LLM annotates on these three dimensions, and test how far the labels hold. Most repositories target data collection and analysis and keep a human in the loop; research repositories cover literature review far more often than data-science ones (33.2% versus 3.2%). Encouragingly, this picture does not hinge on the input: reading skill files instead of summaries changes the methodological and human-in-the-loop shares by less than one point. At the same time, labels depend on how the question is asked: reordering answer options changes the autonomy label for 49.3% of repositories, and separate prompts lower the methodology–autonomy association from Cramér’s V = 0.87 to 0.17. They also vary with the model: without the domain label, GPT-4o-mini, GPT-4o, and GPT-5 put the methodological share of a 600-repository sample at 47.3–58.2%, against 66.2% in the released labels. SciSkillBank thus supports mapping what research communities delegate to agents and shows that a label's question and model belong in its report.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.