
UBITECH’s Energy Digitalization Group (EDG) announces the publication of two research papers that establish rigorous, reproducible benchmarks for evaluating Large Language Model (LLM) agents on real power system engineering workflows. The papers, “PowerAgentBench-SS: A Benchmark for Agentic AI in Power System Steady-State Studies” and “PowerAgentBench-Dyn: A Benchmark for Agentic AI in Power System Dynamic Studies,” are now available on arXiv at https://arxiv.org/abs/2606.18789 and https://arxiv.org/abs/2606.20401 respectively, and will be presented at the North American Power Symposium 2026 (NAPS 2026), taking place October 11-13, 2026 at Michigan Technological University (https://www.mtu.edu/ece/naps-2026/). Both works were developed within UBITECH’s participation in PowerAgent, the open-source Agentic AI community for power systems maintained by Harvard University’s School of Engineering and Applied Sciences (https://poweragent.seas.harvard.edu/), which EDG joined as a member of the core development team. Both papers were authored jointly by researchers from UBITECH’s Energy Digitalization Group, Harvard SEAS, Politecnico di Milano, EnliteAI, and the National Technical University of Athens, with the accompanying benchmark code released publicly through the PowerAgent Community repository at https://github.com/Power-Agent/PowerAgentBench.
The motivation behind the two papers is grounded in a gap that the authors identify across the current benchmarking landscape: existing evaluations of AI in power systems tend to score numerical solvers, prediction models, or sequential controllers in isolation, rather than testing whether an AI agent can actually carry out the kind of end-to-end engineering workflow a power system engineer performs day to day, inspecting a grid case, selecting the right tools, invoking simulators, screening contingencies, proposing admissible mitigations within operational constraints, validating results against physical ground truth, and producing an auditable evidence trail that a human reviewer can trust.
PowerAgentBench-SS addresses this gap for steady-state operation and planning studies. The framework exposes an agent to public case data, explicit action constraints, a defined tool API, and a bounded validation budget, while a hidden evaluator independently recomputes physical validity and scores the agent’s submitted report. The paper formalizes the agent interface, tool contract, and evidence log, and introduces a family of risk-sensitive metrics that go well beyond simple accuracy, including submitted recall, evidence-backed recall, found recall, false-safe penalties, severity regret, residual violation score, action cost, and tool-use efficiency, alongside broader workflow diagnostics. Two task levels anchor the benchmark: an N-1 audit and mitigation level, where the agent must identify violations from a published list of single contingencies and propose bounded corrective controls, and a more demanding N-k search and mitigation level, where the combinatorial contingency space makes exhaustive enumeration infeasible under budget, forcing the agent to allocate its validation calls intelligently through screening, graph reasoning, or adaptive search. The framework is instantiated concretely through a reproducible DC thermal N-2 contingency-search pilot on deterministic IEEE 39-bus operating-point variants, benchmarking scripted baselines against an LLM JSON-command adapter, three locally hosted Ollama agents, and one OpenAI API agent, and the results make a compelling case for why solver-only or answer-only evaluation regimes are insufficient to judge agentic competence in this domain.
PowerAgentBench-Dyn extends the same rigor to power system dynamic studies, a domain the authors describe as particularly promising yet largely unexplored for agentic AI, precisely because dynamic analysis depends less on a single formula and more on a sequence of judgment calls: preparing cases, selecting disturbances, running simulations, diagnosing failed runs, interpreting plots, adjusting parameters within allowed ranges, and building a defensible, evidence-backed conclusion. The benchmark deliberately targets problems that resist reduction to a single optimization or coding task. Its Dynamic Model Quality Review Benchmark mirrors industrial interconnection and reliability workflows, tasking agents with detecting why a submitted renewable resource or large load dynamic model fails quality tests, whether due to unstable controller gains, incorrect ride-through behavior, weak-grid sensitivity, or poor active and reactive power response, and iteratively proposing mitigations. Its companion Dynamic Security Risk Screening Benchmark tests an agent’s ability to use semantic memory and a limited simulation budget to identify, rank, and analyze the most critical short-circuit contingencies from an unseen fault dataset, and to propose and evaluate mitigation measures accordingly. Throughout, the paper carefully separates deterministic task reproducibility, guaranteed by released cases and simulator settings, from the probabilistic reproducibility of agent behavior itself, which is assessed through repeated runs and success statistics rather than single trials.
Taken together, the two benchmarks give the power systems community, and the broader agentic AI research field, a shared, physically grounded yardstick for what “capable” actually means when an autonomous agent is handed real grid engineering responsibility. By exposing constrained action spaces, hidden physical validators, and audit-ready evidence logs, PowerAgentBench-SS and PowerAgentBench-Dyn move the conversation beyond whether an LLM can answer a power systems question correctly, toward whether it can be trusted to operate, one validated decision at a time, inside the safety margins that real grids demand.
“These two benchmarks reflect exactly the kind of contribution our Energy Digitalization Group set out to make when we joined the PowerAgent community,” said Dr. Magda Foti, Head of UBITECH’s Energy Digitalization Group (https://ubitech.eu/edg/). “Power engineers do not work by answering isolated questions, they work through structured, accountable processes with real physical consequences, and any AI agent that hopes to support that work has to be evaluated the same way. With PowerAgentBench-SS and PowerAgentBench-Dyn, we and our partners at Harvard SEAS, Politecnico di Milano, EnliteAI, and NTUA are giving the community reproducible tools to measure that kind of trustworthiness directly, in both steady-state and dynamic studies. We look forward to sharing these results with the power systems community at NAPS 2026.”
Both papers, together with their reproducible benchmark suites, datasets, and evaluation code, are openly available to the research community through the PowerAgent Community GitHub repository at https://github.com/Power-Agent/PowerAgentBench, and the PowerAgent initiative itself can be explored at https://poweragent.seas.harvard.edu/. UBITECH’s Energy Digitalization Group will present both works at the North American Power Symposium 2026, hosted by Michigan Technological University from October 11 to 13, 2026 (https://www.mtu.edu/ece/naps-2026/).

