Skip to content
Developer ToolsOpen sourceArchitecture reviewSelf-hostable
Open Code Review logo

Open Code Review

Open-source AI code review CLI combining deterministic Go file bundling and rule matching with LLM agents for line-precise pull request feedback.

Dhanji Bhagat

Dhanji Bhagat

Founder, Emiote

Managed Cloud

Fully hosted platform. Automated backups and SLA.

Reference Cost

Commercial AI code review at $15-$30/dev/month

Self-Host Path

Private compute. Zero seat taxes; team runs ops.

Reference Cost

$0 local CLI / pay-as-you-go LLM tokens

Open Code Review is an open-source AI code review tool created by Alibaba Group, built in Go. It combines deterministic file selection, sub-agent bundling, and comment reflection with dynamic LLM reasoning to generate line-precise pull request feedback. Designed for local terminals and CI/CD pipelines, it runs via standalone binaries with custom model providers or agent delegation.


1. What It Replaces & Why It Matters

Engineering teams face a growing bottleneck during code review. As AI coding tools accelerate code generation, the volume of pull requests expands rapidly. Traditional peer review struggles to keep pace, while naive AI review prompts introduce severe operational friction.

When developers instruct general-purpose agents like Claude Code, ChatGPT, or Cursor to review a repository diff, three predictable defects appear:

  1. Position Drift: Large language models struggle with relative line arithmetic. Review comments routinely target incorrect line numbers or reference lines outside the modified diff hunk.
  2. Incomplete Coverage: On changesets containing dozens of files, models quietly cut corners. They review the first three or four files in detail and summarize or ignore the remainder.
  3. Alert Fatigue from False Positives: Unconstrained prompts produce floods of superficial commentary regarding style preferences, documentation tone, and trivial renames, obscuring critical concurrency bugs and edge-case errors.

Commercial code review platforms such as CodeRabbit and Qodo attempt to resolve this with proprietary SaaS wrappers. However, they charge between $15 and $30 per developer each month, route private proprietary source code through external vendor clouds, and offer limited flexibility for local command-line workflows.

Open Code Review originated as Alibaba Group’s internal review assistant. Over two years, the underlying engine served tens of thousands of engineers and processed millions of review comments across high-concurrency production systems.

Now open-sourced under the Apache-2.0 license, it replaces closed SaaS subscriptions and fragile prompt wrappers with a standalone Go binary. It enforces hard engineering constraints around file selection, context budgeting, and line coordinate validation before letting models evaluate code semantics.


2. Architecture & Deterministic Hybrid Engine

Open Code Review is written entirely in Go with zero CGO dependencies (CGO_ENABLED=0). It compiles into a single, self-contained binary (ocr) that interacts directly with Git.

flowchart TD
    GitChanges["Local Changes / Pull Request Diff"] --> SelectionModule["Deterministic File Selection (internal/agent/selection.go)\nFilters lockfiles, vendor assets, binaries"]
    SelectionModule --> BundlingModule["Smart File Bundling (internal/agent/grouping.go)\nGroups coupled files (e.g. locale pairs, tests)"]
    
    BundlingModule --> SubAgentPool["Sub-Agent Pool with Isolated Contexts"]
    
    subgraph ReviewUnit ["Isolated Review Unit Execution"]
        SubAgentPool --> RuleEngine["Template-Engine Rule Matcher\nApplies path-specific review guidelines"]
        RuleEngine --> DynamicAgent["Specialized Review Agent (internal/agent)"]
        DynamicAgent <--> ToolCalls["Tuned Toolset (internal/tool):\nfile_read | file_read_diff | code_search"]
        DynamicAgent --> RawComments["Candidate Code Comments (JSON)"]
    end
    
    RawComments --> ReflectionModule["Comment Reflection & Validation (internal/tool/comment_args_repair.go)\nVerifies diff bounds, repairs coordinates, drops false positives"]
    
    ReflectionModule --> OutputPipeline{"Output Router"}
    OutputPipeline --> Terminal["Terminal CLI Output"]
    OutputPipeline --> WebViewer["Local Session Viewer (internal/viewer)"]
    OutputPipeline --> CIPipeline["GitHub Actions / GitLab CI Bot Comments"]

Deterministic Hard Constraints

Open Code Review separates mechanical correctness from subjective evaluation. The system enforces hard engineering constraints where language models are statistically unreliable:

  • Strict File Selection (internal/agent/selection.go): Evaluates file extensions, size limits, and change semantics. It strips vendor libraries, minified bundles, lockfiles, and auto-generated code before allocating token budgets.
  • Smart File Bundling (internal/agent/grouping.go): Groups related files into unified review units. For example, translation files (message_en.properties and message_zh.properties) or coupled interface definitions and unit tests are bundled together. Each unit runs as a sub-agent with an isolated context window, enabling concurrent review execution across large changesets without token spillover.
  • Template-Engine Rule Matching: Instead of injecting hundreds of lines of universal guidelines into every prompt turn, OCR matches review rules against specific file paths and language types using Go templates. This concentrates model attention on language-specific anti-patterns.
  • Comment Positioning and Reflection (internal/tool/comment_args_repair.go): Before any comment reaches the user or CI interface, an independent verification pipeline validates line numbers against actual git diff hunks. If a comment drifts outside the modified range or targets phantom code, the coordinate repair module fixes the offset or discards the hallucination.

Specialized Agent Toolset

Rather than exposing a generic bash shell or broad filesystem tools, the review agent uses five specialized tools tuned through production trace analysis:

Tool NameImplementationFunction
file_readinternal/tool/file_read.goReads full source files within the repository root to evaluate context outside the diff.
file_read_diffinternal/tool/file_read_diff.goInspects exact git diff hunks for a designated file path.
file_findinternal/tool/file_find.goLocates files across the directory tree matching patterns.
code_searchinternal/tool/code_search.goExecutes regex and symbol searches across the codebase to identify call sites and definition points.
code_commentinternal/tool/code_comment.goEmits structured line comments with severity tags, defect classifications, and fix suggestions.

All file operations enforce path containment via pathutil.WithinBase(), blocking directory traversal attacks before and after symlink resolution.


3. Visual Tour & Interface Workflows

Open Code Review operates directly in local development environments, CI runners, and browser interfaces.

Open Code Review Architecture and Feature Highlights

Core Operational Modes

The tool provides three primary execution workflows depending on where review feedback is consumed:

  1. Workspace Diff Review (ocr review): Reviews uncommitted working tree changes, staged index updates, or single commits. It outputs line-level annotations directly into the terminal with syntax-highlighted code diffs.
  2. Branch Range Review (ocr review --from main --to feature): Automatically identifies the merge-base between two git branches, reviewing only code diverged from the upstream trunk. Sessions support persistence and resumption via --resume <session-id>.
  3. Full-File Audits (ocr scan --path <dir>): Audits legacy directories or unfamiliar codebases without requiring git history. It reviews whole files for architectural security defects and code smells.

Delegation Mode

A standout architectural feature is Delegation Mode (ocr delegate preview and ocr delegate rule).

In standard mode, OCR requires its own configured LLM API key. In Delegation Mode, OCR serves as an orchestration engine for host coding agents like Claude Code, Cursor, or Codex.

OCR generates the deterministic file bundles, resolves path-specific review rules, and packages the exact diff payloads. The host coding agent then evaluates the review using its existing session tokens. This eliminates the need for separate API credentials in corporate environments.

Local Session Viewer

Running ocr viewer starts an embedded HTTP server providing a graphical interface for reviewing session history.

Developers can filter comments by severity, mark false positives as ignored, and hide completed items while working through findings. To prevent security vulnerabilities on developer machines, the viewer binds to the loopback interface by default and uses internal/viewer/hostguard.go to reject DNS rebinding attempts from non-local host headers.


4. Total Cost of Ownership (TCO)

Operating an automated code review system involves license fees, compute overhead, and model inference tokens. Open Code Review provides massive cost advantages over both commercial SaaS products and unstructured agent prompts.

DimensionCommercial Review SaaS (CodeRabbit, Qodo)Unstructured Agent Prompts (Claude Code / Cursor)Open Code Review (CLI + OSS Engine)
License Cost$15 to $30 per developer / monthIncluded in assistant seat subscription$0 (Apache-2.0 open source)
Token EfficiencyProprietary cloud cachingLow (repeats full file contents every turn)High (consumes ~1/9 tokens of raw agents)
API Token BillsBundled into seat license$0.15 to $0.60 per pull request review$0.02 to $0.08 per review (direct API pricing)
InfrastructureVendor-managed multi-tenant cloudLocal assistant desktop runtimeLocal Go binary (<50MB) / zero-dependency CI
Data PrivacyCode sent to third-party vendor cloudsModel provider privacy policyDirect connection to private or VPC models
Delegation ModeNot supportedNot applicableSupported ($0 additional API configuration)
Annual TCO (10 Devs)$1,800 to $3,600 / yearVariable token consumption$0 software + direct pay-as-you-go tokens

Benchmark Evidence: The AACR-Bench Dataset

Alibaba open-sourced its evaluation benchmark, AACR-Bench, hosted publicly on Hugging Face (Alibaba-Aone/aacr-bench). The benchmark comprises:

  • 50 popular open-source repositories
  • 200 real pull requests across 10 programming languages
  • 1,505 ground-truth issues annotated and cross-validated by 80+ senior engineers

When evaluated against general-purpose agents using identical foundation models, Open Code Review achieved significantly higher Precision and F1 scores while consuming approximately one-ninth (1/9) of the input tokens.

By pre-filtering noise and isolating review files into bounded sub-agent units, it eliminates redundant context passing, directly cutting API inference costs by up to 88%.


5. The Bad: What to Know Before Adopting

Before deploying Open Code Review across production teams, engineers should understand several concrete operational trade-offs:

  1. Deliberate Low Recall Policy: Open Code Review explicitly favors Precision over Recall. The default filtering and reflection modules are tuned aggressively to eliminate developer alert fatigue. Consequently, the tool will intentionally overlook subjective stylistic choices, minor documentation phrasing, and speculative edge cases. Teams seeking an exhaustive linter that catches every pedantic violation will find its output conservative.
  2. Comment Reflection Fragility on Small Models: The reflection and line coordinate repair pipeline relies on strict JSON schema compliance and spatial reasoning. When running against smaller local models (such as sub-14B parameter models like Qwen-2.5-Coder-7B), the reflection module frequently fails to resolve line number repairs, resulting in discarded comments. The tool functions reliably only when paired with frontier reasoning models (Claude 3.5/3.7 Sonnet, GPT-4o, DeepSeek-V3, or Qwen-2.5-Coder-32B).
  3. Git Merge-Base Requirements in Shallow CI Environments: In range-review mode (ocr review --from main --to feature), the engine relies on git merge-base to detect the divergence commit. In CI environments where repositories are cloned with shallow history (fetch-depth: 1), the merge-base command fails unless the CI pipeline explicitly fetches the upstream base branch.
  4. Local Web Viewer Host Restrictions: The built-in session viewer enforces strict host header checking via hostguard.go. When running OCR inside remote development containers (e.g. GitHub Codespaces, remote SSH sessions, or Docker workspaces), accessing the viewer through port forwarding requires configuring OCR_VIEWER_ALLOWED_HOSTS to avoid HTTP 403 Forbidden errors.

6. Quickstart & Deployment

Open Code Review distributes as a pre-compiled Go binary, an npm global package, or a GitHub Action.

Global CLI Installation

# Install via npm package manager
npm install -g @alibaba-group/open-code-review

# Verify installation
ocr --version

Alternatively, download the standalone binary directly for Linux, macOS, or Windows from the project’s GitHub Releases page.

Configure Model Provider

# Interactive provider setup (supports OpenAI, Anthropic, Qwen, DeepSeek, Ollama)
ocr config provider

# Select active model
ocr config model

Run Review

# Review uncommitted changes in current repository
ocr review

# Review pull request changes between branches
ocr review --from main --to feature-branch

# Export structured review findings to JSON
ocr review --format json --output review-results.json

GitHub Actions CI Integration

To automate reviews on pull requests, add Open Code Review to your workflow:

name: AI Code Review
on:
  pull_request:
    types: [opened, synchronize]

jobs:
  review:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0 # Full history required for git merge-base

      - name: Run Open Code Review
        uses: alibaba/open-code-review@main
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
        with:
          provider: "openai"
          model: "gpt-4o"

7. Recommendation & ReframeHub Insight

Who Should Use This

  • Engineering Teams with High PR Velocity: Organizations seeking to accelerate code review cycles without overwhelming senior developers with initial sanity checking.
  • Privacy-Sensitive Organizations: Teams that cannot transmit proprietary code to third-party SaaS vendors and must run code review through private model endpoints or self-hosted VPCs.
  • Developers Using AI Coding Assistants: Engineers using Claude Code, Cursor, or Windsurf who want deterministic file bundling and rule resolution via Delegation Mode.

Who Should Avoid This

  • Teams Seeking Full Style Enforcement: Projects that require pedantic linting and formatting feedback; standard static analysis tools (ESLint, golangci-lint, Ruff) handle syntax rules far more reliably.
  • Environments Restricted to Low-Resource Local Models: Setups attempting to run review agents on sub-14B parameter models without cloud API access.

ReframeHub Insight: Hard Constraints Protect Agent Focus

The core engineering lesson of Open Code Review is that language models should never be tasked with deterministic bookkeeping.

When developers build review bots, the common temptation is to write a comprehensive prompt and feed it raw git diffs. This approach forces the neural network to calculate line offsets, filter vendor directories, and balance token quotas. Language models are probabilistic pattern engines; asking them to perform coordinate arithmetic guarantees position drift and silent truncation.

Open Code Review succeeds because it sandwiches the language model between two deterministic engineering layers:

  1. Upstream Pre-Processing: Go routines inspect Git trees, filter noise, bundle coupled files, and calculate exact token budgets before calling the LLM.
  2. Downstream Post-Processing: Dedicated reflection algorithms inspect candidate comments, verify diff ranges, and drop hallucinations before output.

The language model is restricted to the single task where it holds a clear advantage: evaluating semantic code logic. By eliminating mechanical overhead, Open Code Review reduces token consumption by 88% while producing reviews that developers actually trust.