Open Code Review
Open-source AI code review CLI combining deterministic Go file bundling and rule matching with LLM agents for line-precise pull request feedback.

Dhanji Bhagat
Founder, Emiote
Fully hosted platform. Automated backups and SLA.
Commercial AI code review at $15-$30/dev/month
Private compute. Zero seat taxes; team runs ops.
$0 local CLI / pay-as-you-go LLM tokens
Open Code Review is an open-source AI code review tool created by Alibaba Group, built in Go. It combines deterministic file selection, sub-agent bundling, and comment reflection with dynamic LLM reasoning to generate line-precise pull request feedback. Designed for local terminals and CI/CD pipelines, it runs via standalone binaries with custom model providers or agent delegation.
1. What It Replaces & Why It Matters
Engineering teams face a growing bottleneck during code review. As AI coding tools accelerate code generation, the volume of pull requests expands rapidly. Traditional peer review struggles to keep pace, while naive AI review prompts introduce severe operational friction.
When developers instruct general-purpose agents like Claude Code, ChatGPT, or Cursor to review a repository diff, three predictable defects appear:
- Position Drift: Large language models struggle with relative line arithmetic. Review comments routinely target incorrect line numbers or reference lines outside the modified diff hunk.
- Incomplete Coverage: On changesets containing dozens of files, models quietly cut corners. They review the first three or four files in detail and summarize or ignore the remainder.
- Alert Fatigue from False Positives: Unconstrained prompts produce floods of superficial commentary regarding style preferences, documentation tone, and trivial renames, obscuring critical concurrency bugs and edge-case errors.
Commercial code review platforms such as CodeRabbit and Qodo attempt to resolve this with proprietary SaaS wrappers. However, they charge between $15 and $30 per developer each month, route private proprietary source code through external vendor clouds, and offer limited flexibility for local command-line workflows.
Open Code Review originated as Alibaba Group’s internal review assistant. Over two years, the underlying engine served tens of thousands of engineers and processed millions of review comments across high-concurrency production systems.
Now open-sourced under the Apache-2.0 license, it replaces closed SaaS subscriptions and fragile prompt wrappers with a standalone Go binary. It enforces hard engineering constraints around file selection, context budgeting, and line coordinate validation before letting models evaluate code semantics.
2. Architecture & Deterministic Hybrid Engine
Open Code Review is written entirely in Go with zero CGO dependencies (CGO_ENABLED=0). It compiles into a single, self-contained binary (ocr) that interacts directly with Git.
flowchart TD
GitChanges["Local Changes / Pull Request Diff"] --> SelectionModule["Deterministic File Selection (internal/agent/selection.go)\nFilters lockfiles, vendor assets, binaries"]
SelectionModule --> BundlingModule["Smart File Bundling (internal/agent/grouping.go)\nGroups coupled files (e.g. locale pairs, tests)"]
BundlingModule --> SubAgentPool["Sub-Agent Pool with Isolated Contexts"]
subgraph ReviewUnit ["Isolated Review Unit Execution"]
SubAgentPool --> RuleEngine["Template-Engine Rule Matcher\nApplies path-specific review guidelines"]
RuleEngine --> DynamicAgent["Specialized Review Agent (internal/agent)"]
DynamicAgent <--> ToolCalls["Tuned Toolset (internal/tool):\nfile_read | file_read_diff | code_search"]
DynamicAgent --> RawComments["Candidate Code Comments (JSON)"]
end
RawComments --> ReflectionModule["Comment Reflection & Validation (internal/tool/comment_args_repair.go)\nVerifies diff bounds, repairs coordinates, drops false positives"]
ReflectionModule --> OutputPipeline{"Output Router"}
OutputPipeline --> Terminal["Terminal CLI Output"]
OutputPipeline --> WebViewer["Local Session Viewer (internal/viewer)"]
OutputPipeline --> CIPipeline["GitHub Actions / GitLab CI Bot Comments"]
Deterministic Hard Constraints
Open Code Review separates mechanical correctness from subjective evaluation. The system enforces hard engineering constraints where language models are statistically unreliable:
- Strict File Selection (
internal/agent/selection.go): Evaluates file extensions, size limits, and change semantics. It strips vendor libraries, minified bundles, lockfiles, and auto-generated code before allocating token budgets. - Smart File Bundling (
internal/agent/grouping.go): Groups related files into unified review units. For example, translation files (message_en.propertiesandmessage_zh.properties) or coupled interface definitions and unit tests are bundled together. Each unit runs as a sub-agent with an isolated context window, enabling concurrent review execution across large changesets without token spillover. - Template-Engine Rule Matching: Instead of injecting hundreds of lines of universal guidelines into every prompt turn, OCR matches review rules against specific file paths and language types using Go templates. This concentrates model attention on language-specific anti-patterns.
- Comment Positioning and Reflection (
internal/tool/comment_args_repair.go): Before any comment reaches the user or CI interface, an independent verification pipeline validates line numbers against actual git diff hunks. If a comment drifts outside the modified range or targets phantom code, the coordinate repair module fixes the offset or discards the hallucination.
Specialized Agent Toolset
Rather than exposing a generic bash shell or broad filesystem tools, the review agent uses five specialized tools tuned through production trace analysis:
| Tool Name | Implementation | Function |
|---|---|---|
file_read | internal/tool/file_read.go | Reads full source files within the repository root to evaluate context outside the diff. |
file_read_diff | internal/tool/file_read_diff.go | Inspects exact git diff hunks for a designated file path. |
file_find | internal/tool/file_find.go | Locates files across the directory tree matching patterns. |
code_search | internal/tool/code_search.go | Executes regex and symbol searches across the codebase to identify call sites and definition points. |
code_comment | internal/tool/code_comment.go | Emits structured line comments with severity tags, defect classifications, and fix suggestions. |
All file operations enforce path containment via pathutil.WithinBase(), blocking directory traversal attacks before and after symlink resolution.
3. Visual Tour & Interface Workflows
Open Code Review operates directly in local development environments, CI runners, and browser interfaces.

Core Operational Modes
The tool provides three primary execution workflows depending on where review feedback is consumed:
- Workspace Diff Review (
ocr review): Reviews uncommitted working tree changes, staged index updates, or single commits. It outputs line-level annotations directly into the terminal with syntax-highlighted code diffs. - Branch Range Review (
ocr review --from main --to feature): Automatically identifies the merge-base between two git branches, reviewing only code diverged from the upstream trunk. Sessions support persistence and resumption via--resume <session-id>. - Full-File Audits (
ocr scan --path <dir>): Audits legacy directories or unfamiliar codebases without requiring git history. It reviews whole files for architectural security defects and code smells.
Delegation Mode
A standout architectural feature is Delegation Mode (ocr delegate preview and ocr delegate rule).
In standard mode, OCR requires its own configured LLM API key. In Delegation Mode, OCR serves as an orchestration engine for host coding agents like Claude Code, Cursor, or Codex.
OCR generates the deterministic file bundles, resolves path-specific review rules, and packages the exact diff payloads. The host coding agent then evaluates the review using its existing session tokens. This eliminates the need for separate API credentials in corporate environments.
Local Session Viewer
Running ocr viewer starts an embedded HTTP server providing a graphical interface for reviewing session history.
Developers can filter comments by severity, mark false positives as ignored, and hide completed items while working through findings. To prevent security vulnerabilities on developer machines, the viewer binds to the loopback interface by default and uses internal/viewer/hostguard.go to reject DNS rebinding attempts from non-local host headers.
4. Total Cost of Ownership (TCO)
Operating an automated code review system involves license fees, compute overhead, and model inference tokens. Open Code Review provides massive cost advantages over both commercial SaaS products and unstructured agent prompts.
| Dimension | Commercial Review SaaS (CodeRabbit, Qodo) | Unstructured Agent Prompts (Claude Code / Cursor) | Open Code Review (CLI + OSS Engine) |
|---|---|---|---|
| License Cost | $15 to $30 per developer / month | Included in assistant seat subscription | $0 (Apache-2.0 open source) |
| Token Efficiency | Proprietary cloud caching | Low (repeats full file contents every turn) | High (consumes ~1/9 tokens of raw agents) |
| API Token Bills | Bundled into seat license | $0.15 to $0.60 per pull request review | $0.02 to $0.08 per review (direct API pricing) |
| Infrastructure | Vendor-managed multi-tenant cloud | Local assistant desktop runtime | Local Go binary (<50MB) / zero-dependency CI |
| Data Privacy | Code sent to third-party vendor clouds | Model provider privacy policy | Direct connection to private or VPC models |
| Delegation Mode | Not supported | Not applicable | Supported ($0 additional API configuration) |
| Annual TCO (10 Devs) | $1,800 to $3,600 / year | Variable token consumption | $0 software + direct pay-as-you-go tokens |
Benchmark Evidence: The AACR-Bench Dataset
Alibaba open-sourced its evaluation benchmark, AACR-Bench, hosted publicly on Hugging Face (Alibaba-Aone/aacr-bench). The benchmark comprises:
- 50 popular open-source repositories
- 200 real pull requests across 10 programming languages
- 1,505 ground-truth issues annotated and cross-validated by 80+ senior engineers
When evaluated against general-purpose agents using identical foundation models, Open Code Review achieved significantly higher Precision and F1 scores while consuming approximately one-ninth (1/9) of the input tokens.
By pre-filtering noise and isolating review files into bounded sub-agent units, it eliminates redundant context passing, directly cutting API inference costs by up to 88%.
5. The Bad: What to Know Before Adopting
Before deploying Open Code Review across production teams, engineers should understand several concrete operational trade-offs:
- Deliberate Low Recall Policy: Open Code Review explicitly favors Precision over Recall. The default filtering and reflection modules are tuned aggressively to eliminate developer alert fatigue. Consequently, the tool will intentionally overlook subjective stylistic choices, minor documentation phrasing, and speculative edge cases. Teams seeking an exhaustive linter that catches every pedantic violation will find its output conservative.
- Comment Reflection Fragility on Small Models: The reflection and line coordinate repair pipeline relies on strict JSON schema compliance and spatial reasoning. When running against smaller local models (such as sub-14B parameter models like Qwen-2.5-Coder-7B), the reflection module frequently fails to resolve line number repairs, resulting in discarded comments. The tool functions reliably only when paired with frontier reasoning models (Claude 3.5/3.7 Sonnet, GPT-4o, DeepSeek-V3, or Qwen-2.5-Coder-32B).
- Git Merge-Base Requirements in Shallow CI Environments: In range-review mode (
ocr review --from main --to feature), the engine relies ongit merge-baseto detect the divergence commit. In CI environments where repositories are cloned with shallow history (fetch-depth: 1), the merge-base command fails unless the CI pipeline explicitly fetches the upstream base branch. - Local Web Viewer Host Restrictions: The built-in session viewer enforces strict host header checking via
hostguard.go. When running OCR inside remote development containers (e.g. GitHub Codespaces, remote SSH sessions, or Docker workspaces), accessing the viewer through port forwarding requires configuringOCR_VIEWER_ALLOWED_HOSTSto avoid HTTP 403 Forbidden errors.
6. Quickstart & Deployment
Open Code Review distributes as a pre-compiled Go binary, an npm global package, or a GitHub Action.
Global CLI Installation
# Install via npm package manager
npm install -g @alibaba-group/open-code-review
# Verify installation
ocr --version
Alternatively, download the standalone binary directly for Linux, macOS, or Windows from the project’s GitHub Releases page.
Configure Model Provider
# Interactive provider setup (supports OpenAI, Anthropic, Qwen, DeepSeek, Ollama)
ocr config provider
# Select active model
ocr config model
Run Review
# Review uncommitted changes in current repository
ocr review
# Review pull request changes between branches
ocr review --from main --to feature-branch
# Export structured review findings to JSON
ocr review --format json --output review-results.json
GitHub Actions CI Integration
To automate reviews on pull requests, add Open Code Review to your workflow:
name: AI Code Review
on:
pull_request:
types: [opened, synchronize]
jobs:
review:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0 # Full history required for git merge-base
- name: Run Open Code Review
uses: alibaba/open-code-review@main
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
with:
provider: "openai"
model: "gpt-4o"
7. Recommendation & ReframeHub Insight
Who Should Use This
- Engineering Teams with High PR Velocity: Organizations seeking to accelerate code review cycles without overwhelming senior developers with initial sanity checking.
- Privacy-Sensitive Organizations: Teams that cannot transmit proprietary code to third-party SaaS vendors and must run code review through private model endpoints or self-hosted VPCs.
- Developers Using AI Coding Assistants: Engineers using Claude Code, Cursor, or Windsurf who want deterministic file bundling and rule resolution via Delegation Mode.
Who Should Avoid This
- Teams Seeking Full Style Enforcement: Projects that require pedantic linting and formatting feedback; standard static analysis tools (ESLint, golangci-lint, Ruff) handle syntax rules far more reliably.
- Environments Restricted to Low-Resource Local Models: Setups attempting to run review agents on sub-14B parameter models without cloud API access.
ReframeHub Insight: Hard Constraints Protect Agent Focus
The core engineering lesson of Open Code Review is that language models should never be tasked with deterministic bookkeeping.
When developers build review bots, the common temptation is to write a comprehensive prompt and feed it raw git diffs. This approach forces the neural network to calculate line offsets, filter vendor directories, and balance token quotas. Language models are probabilistic pattern engines; asking them to perform coordinate arithmetic guarantees position drift and silent truncation.
Open Code Review succeeds because it sandwiches the language model between two deterministic engineering layers:
- Upstream Pre-Processing: Go routines inspect Git trees, filter noise, bundle coupled files, and calculate exact token budgets before calling the LLM.
- Downstream Post-Processing: Dedicated reflection algorithms inspect candidate comments, verify diff ranges, and drop hallucinations before output.
The language model is restricted to the single task where it holds a clear advantage: evaluating semantic code logic. By eliminating mechanical overhead, Open Code Review reduces token consumption by 88% while producing reviews that developers actually trust.
Automating code review in CI or local agent loops?
Reframe ($199) audits your code review pipelines, LLM token efficiency, comment reflection accuracy, and developer alert fatigue.
Fixed $199 fee · 100% vendor-neutral review · 3-day delivery guarantee
