Back to The Proceedings

Research papers · 4 min read

GraphZero Agent: Monte Carlo Graph Search as Reasoning-as-a-Service

Today's agents reason along a single path and rarely look back. This paper adapts AlphaZero-style search to general agentic reasoning — over a graph rather than a tree — and offers it as a service any agent framework can call.

KR
Karthik Rajkumar14 March 2026 · 4 min read
  • Published in: International Journal of Scientific Research in Engineering and Management (IJSREM), Vol. 10, No. 3, March 2026
  • Authors: Karthik Rajkumar Kannan, Bipin Chandra
  • DOI: 10.55041/IJSREM57596
  • Review: double-blind peer review, two rounds
  • Licence: open access under CC BY 4.0

Read the full paper (PDF) · View on IJSREM

#Abstract

Current agentic AI systems, predominantly based on the Reason-Act (ReAct) paradigm, execute single-trajectory reasoning with limited backtracking, no principled exploration, and no reuse of repeated subproblems. Empirical analyses reveal brittleness to prompt perturbations and catastrophic performance degradation on tasks exceeding 15–120 steps. We propose GraphZero Agent, a reasoning architecture that adapts AlphaZero's policy/value-guided Monte Carlo Tree Search (MCTS) to Monte Carlo Graph Search (MCGS) over general agentic belief-states. GraphZero Agent generalizes the search space from trees to Directed Acyclic Graphs, enabling transposition merging of semantically equivalent reasoning states via dense embedding similarity and Approximate Nearest Neighbor search. Intermediate states are scored by Process Reward Models (PRMs) trained without human annotation through a self-rewarding rollout pipeline that produces Process Preference Models (PPMs) from pairwise trajectory comparisons. Beyond single-session inference, a dual-loop Empirical-MCTS architecture provides continuous non-parametric learning: a local loop refines the LLM's generation policy within a search session via contrastive meta- prompting (PE-EMP), while a global loop distills winning trajectories into reusable heuristics across sessions. The system is deployed as Reasoning-as-a- Service (RaaS) via Model Context Protocol (MCP) or OpenAPI, enabling any agentic orchestration framework to invoke deep deliberative search with independent scaling and LLM hot-swapping.

#The paper in plain terms

#The problem: agents that think in a straight line

Most agents today follow the ReAct pattern: think a step, act, observe, think the next step. It is simple and it works for short tasks. But it commits to a single line of reasoning. There is little backtracking when a step turns out to be wrong, no principled way to explore alternatives, and no memory that the same sub-problem was already solved a few steps earlier.

The consequences show up on long tasks. The paper points to analyses showing that these agents are brittle to small changes in the prompt and degrade sharply once a task runs beyond roughly 15 to 120 steps.

#The idea: search, the way AlphaZero does

AlphaZero mastered games like Go and chess by not committing to the first move that looks good. It runs a Monte Carlo Tree Search: it explores many possible continuations, guided by a policy that suggests promising moves and a value estimate that judges positions, and spends its effort on the branches that matter.

GraphZero Agent brings that deliberate, search-based reasoning to general agentic tasks. Instead of positions on a board, the states it searches over are the agent's belief-states — what it knows and has concluded so far.

#From trees to graphs

A search tree treats every path as distinct, even when two different routes arrive at the same place. In reasoning that happens constantly: two chains of thought can reach what is, in substance, the same intermediate conclusion.

GraphZero Agent searches a directed acyclic graph instead. When two reasoning states are semantically equivalent — detected by comparing dense embeddings with approximate nearest-neighbour search — they are merged into one node. Work done on a sub-problem is shared rather than repeated.

#Scoring the steps, not just the answer

Search needs a way to judge intermediate states, not only final answers. GraphZero Agent uses process reward models to score each step. Crucially, these are trained without human annotation: a self-rewarding rollout pipeline compares pairs of trajectories and learns process preference models from those comparisons.

#Learning across sessions

The architecture also improves as it runs, without retraining the underlying model. Two loops do this:

  • A local loop refines how the language model generates candidate steps within a single search session, using contrastive meta-prompting.
  • A global loop distils the trajectories that won into reusable heuristics that carry over to future sessions.

#Reasoning as a service

Finally, GraphZero Agent is packaged as a service rather than a framework. Any agent orchestration system can call it through the Model Context Protocol (MCP) or an OpenAPI interface when it needs deep, deliberate reasoning. The search scales independently of the agent calling it, and the underlying language model can be swapped without changing the rest of the system.

#Cite this paper

IEEEText
K. R. Kannan and B. Chandra, "GraphZero Agent: Monte Carlo Graph Search as Reasoning-as-a-Service for General Agentic Systems," International Journal of Scientific Research in Engineering and Management, vol. 10, no. 3, Mar. 2026, doi: 10.55041/IJSREM57596.
BibTeXText
@article{kannan2026graphzero,
  title   = {GraphZero Agent: Monte Carlo Graph Search as Reasoning-as-a-Service
             for General Agentic Systems},
  author  = {Kannan, Karthik Rajkumar and Chandra, Bipin},
  journal = {International Journal of Scientific Research in Engineering and Management},
  volume  = {10},
  number  = {3},
  year    = {2026},
  month   = mar,
  doi     = {10.55041/IJSREM57596}
}

Found this useful? Pass it on.