Skip to content
Developer ToolsOpen sourceArchitecture reviewSelf-hostable
X Algorithm logo

X Algorithm

Deep dive into xAI's open-source Rust Home Mixer, Phoenix recommendation models, candidate retrieval, ranking, and visibility filtering pipeline.

Dhanji Bhagat

Dhanji Bhagat

Founder, Emiote

Managed Cloud

Fully hosted platform. Automated backups and SLA.

Reference Cost

Not applicable; production-scale infrastructure costs are not published by the repository

Self-Host Path

Private compute. Zero seat taxes; team runs ops.

Reference Cost

Open source; the repository includes a runnable nano Phoenix training/retrieval path for local experimentation

The X Algorithm (xai-org/x-algorithm) is the open-source codebase for the For You feed recommendation system on X. The repository exposes the Rust-based Home Mixer orchestration and feed pipeline, candidate retrieval through Thunder, Phoenix, and SimClusters, learned ranking models, visibility filtering, and supporting training/reference infrastructure.


1. Why the X Algorithm Matters

Recommendation systems are often treated as black boxes.

X has now open-sourced a substantial portion of the architecture behind its For You feed, making it possible to inspect how candidate generation, filtering, machine-learned ranking, selection, and visibility decisions fit together.

The important architectural idea is not that every decision is made by one neural network.

Instead, the system combines:

  1. Learned relevance models for predicting viewer behavior.
  2. Fast candidate sources for posts from followed and non-followed accounts.
  3. Deterministic filters and business logic around the ranking model.
  4. Visibility and safety systems that can determine whether a post is allowed to appear at all.
  5. Configuration and experimentation that control how the pipeline behaves.
flowchart TD
    subgraph Ingestion["1. Context & Query Hydration"]
        Context["Viewer Context<br/>(Action history, social graph, blocks/mutes, topics)"]
    end

    subgraph Retrieval["2. Candidate Retrieval"]
        Thunder["In-Network: Thunder<br/>(Followed accounts in-memory store)"]
        PhoenixRetrieval["Out-of-Network: Phoenix<br/>(ANN vector embedding index)"]
        SimClusters["Out-of-Network: SimClusters<br/>(Community cluster representations)"]
    end

    subgraph Processing["3. Candidate Hydration & Pre-Filters"]
        Hydration["Candidate Hydration<br/>(Post text, media, author features, subscriptions)"]
        PreFilters["Pre-Scoring Filters<br/>(Deduplication, age thresholds, mute/block lists)"]
    end

    subgraph Scoring["4. Neural Ranking & Scoring"]
        Ranker["Phoenix Neural Ranker<br/>(Multi-action transformer predictions)"]
        Scorer["RankingScorer<br/>(Weighted sum of action probabilities)"]
        Adjustments["Ranking Adjustments<br/>(Author diversity decay, out-of-network boosts)"]
    end

    subgraph Delivery["5. Visibility Filtering & Blending"]
        Visibility{"Visibility Filtering<br/>(ALLOW / INTERSTITIAL / DROP)"}
        Blending["Blending Pipeline<br/>(Interleaves Ads, Prompts, Who to Follow)"]
        Feed["Final For You Feed Delivery"]
    end

    Context --> Thunder & PhoenixRetrieval & SimClusters
    Thunder & PhoenixRetrieval & SimClusters --> Hydration
    Hydration --> PreFilters
    PreFilters --> Ranker
    Ranker --> Scorer
    Scorer --> Adjustments
    Adjustments --> Visibility
    Visibility -->|ALLOW| Blending
    Blending --> Feed

2. Core Architecture & Request Lifecycle

When a For You request is processed, the repository describes a pipeline roughly like this:

Query hydration → candidate retrieval → candidate hydration → pre-scoring filters → scoring → selection → post-selection filtering → blending

1. Query Hydration

Home Mixer gathers viewer context, including:

  • Recent user action history.
  • Following relationships.
  • Blocks and mutes.
  • Muted keywords.
  • Followed topics.
  • Posts already seen or served.

This context is used both for candidate generation and for personalized ranking.

2. Candidate Retrieval

The repository separates candidate sources into in-network and out-of-network paths.

In-network

Thunder provides recent posts from accounts the viewer follows.

Out-of-network

Phoenix retrieval and SimClusters provide candidates from outside the viewer’s direct network.

The important point is that X does not simply rank the posts from accounts you follow. The For You feed intentionally combines content from your existing network with content discovered outside it.

3. Candidate Hydration

Candidates are enriched with information needed by later stages, such as:

  • Post text and media.
  • Author information and labels.
  • Quoted-post information.
  • Language.
  • Engagement information.
  • Subscription/access information.

4. Pre-Scoring Filters

Before ranking, the pipeline can remove candidates for reasons including:

  • Duplicate posts across sources.
  • Posts older than the configured age threshold.
  • The viewer’s own posts.
  • Blocked or muted accounts.
  • Muted keywords.
  • Posts already seen or served.
  • Inaccessible subscriber-only posts.

3. Phoenix: Retrieval and Ranking Models

One of the easiest ways to misunderstand the repository is to describe Phoenix as simply a conventional “two-tower model.”

The current codebase contains both retrieval and ranking paths, and the ranking model uses a transformer-style architecture over user/history and candidate representations.

The Phoenix ranking model predicts multiple engagement outcomes simultaneously rather than producing one unexplained relevance number.

The repository’s model documentation describes attention relationships between:

  • User/history representations.
  • Candidate representations.
  • Candidate-to-candidate interactions.

This makes Phoenix better described as a multi-action transformer-based recommendation model with separate retrieval and ranking paths, rather than reducing the entire system to cosine similarity between one user vector and one post vector.

flowchart TD
    subgraph Inputs["Input Representations"]
        UserRepr["Viewer Context & History<br/>(Recent engagements, clicks, mutes)"]
        PostRepr["Candidate Post Features<br/>(Post text, author signals, media)"]
    end

    subgraph RetrievalPath["Path A: Phoenix Retrieval (Two-Tower ANN)"]
        UserTower["User Embedding Tower"]
        PostTower["Candidate Embedding Tower"]
        ANN["Approximate Nearest Neighbor (ANN)<br/>Fast sub-millisecond retrieval from millions"]
    end

    subgraph RankingPath["Path B: Phoenix Neural Ranker (Multi-Action Transformer)"]
        CrossAttn["Cross-Attention & Transformer Layers<br/>(Viewer-to-candidate & candidate-to-candidate interactions)"]

        subgraph MultiTask["Simultaneous Action Heads"]
            Fav["P(Favorite)"]
            Rep["P(Reply)"]
            Ret["P(Repost)"]
            Clk["P(Click / Dwell)"]
            Neg["P(Negative Signal)"]
        end

        RankingScorer["RankingScorer<br/>Score = ∑ (wₐ · P(actionₐ))"]
    end

    UserRepr --> UserTower
    PostRepr --> PostTower
    UserTower & PostTower --> ANN
    ANN -->|"Retrieved Candidate Pool"| CrossAttn

    UserRepr & PostRepr --> CrossAttn
    CrossAttn --> Fav & Rep & Ret & Clk & Neg
    Fav & Rep & Ret & Clk & Neg --> RankingScorer

What the model is trying to predict

The ranking model contains multiple engagement targets, including actions such as:

  • Favorite
  • Reply
  • Repost
  • Click
  • Share
  • Profile click
  • Dwell/watch-time related signals
  • Negative feedback signals

The exact set and configuration are represented in the repository’s model and ranking code.


4. Scoring, Ranking & Visibility Filtering

This is one of the most interesting parts of the source code.

Multi-Action Prediction

Phoenix produces predictions for multiple possible actions.

The RankingScorer then combines predicted action values using configurable weights.

Conceptually:

Score = ∑ [ wₐ · P(actionₐ | viewer, post) ]

with additional ranking adjustments and gates applied by the pipeline.

An Important Misconception

The repository explicitly warns against interpreting ranking weights as raw engagement-count equivalents.

For example, it would be wrong to conclude:

“One report cancels out hundreds of likes.”

The weights apply to the predicted probability/value of an action, not raw counts of actions.

This distinction matters because those predictions are personalized to the viewer.

A Like weight therefore does not mean:

1 Like = X ranking points.

It means the model’s predicted likelihood/value of that viewer taking the action contributes to the candidate’s score according to the configured weight.

Ranking Adjustments

The ranking code also contains mechanisms such as:

  • Repeated-author decay.
  • Out-of-network adjustments.
  • New/unexplored-post signals.
  • Negative-feedback signals.
  • Special boosts under particular conditions.

These mechanisms sit around the learned predictions rather than replacing them.

Selection

After scoring, the system selects the highest-ranked candidates.

Visibility Filtering

Ranking is not the same thing as permission to show a post.

The repository contains a separate visibility-filtering system that can determine whether a candidate should:

  • ALLOW - appear normally.
  • INTERSTITIAL - appear behind an additional warning/interstitial.
  • DROP - not appear.

This means a highly ranked post can still be removed by downstream visibility or safety logic.


5. The Full For You Feed Is More Than Ranking

The repository describes two major paths:

Post Pipeline

This path handles finding, ranking, filtering, and selecting posts.

Blending Pipeline

The final For You experience also includes items that are not simply ranked posts, such as:

  • Ads
  • Who to Follow recommendations
  • Prompts
  • Other feed-level items

A blending stage interleaves these with ranked posts.

This is an important architectural distinction:

The recommendation model is only one component of the complete feed system.


6. The Labeling & Safety Path

Visibility decisions can depend on labels generated outside the immediate request-time ranking flow.

The repository includes systems for:

  • Text/content classification.
  • Image and video analysis.
  • Account-level signals.
  • Rule-based labeling.
  • Abuse enforcement.
  • Safety-related aggregation.

These labels can be stored and later consumed by visibility filtering.

This creates a separate path:

Content/account analysis → labels → storage → visibility filtering

rather than requiring every safety decision to be computed from scratch during every feed request.


7. Training & Local Experimentation

The repository includes a runnable nano Phoenix reference path for experimentation.

The documented quickstart can:

  1. Generate synthetic data.
  2. Train nano ranking and retrieval models.
  3. Write checkpoints.
  4. Resume training.
  5. Serve the models.
  6. Send a retrieve → rank request.

Important limitation:

This is a reference/proof-of-concept path, not a reproduction of X’s production-scale recommendation infrastructure.

The repository explicitly states that production data, production checkpoints, orchestration, and production scale are not included.

A simplified starting point is:

git clone https://github.com/xai-org/x-algorithm.git
cd x-algorithm

# Follow the repository's Phoenix QUICKSTART.md
# to generate synthetic data, train the nano models,
# serve them, and run a retrieve -> rank request.

The Bad

  • Production scale is not reproducible locally without proprietary distributed training clusters.
  • Safety and visibility filtering operate as separate pipelines; high ranking scores do not guarantee post display.
  • Synthetic data generation in the quickstart serves as an architectural reference rather than production infrastructure.

Visual Tour / Workflow

xAI Recommendation Algorithm Repository and Architecture Pipeline

  1. Repository Layout: Rust-based Home Mixer coordinates retrieval, scoring, and blending across modular submodules.
  2. Hydration Phase: User actions, blocks, mutes, and follow graphs hydrate candidate post queries in parallel.
  3. Scoring & Selection: Two-tower Phoenix models compute interaction probabilities before downstream visibility filters apply policy rules.

Quickstart / Deployment

Clone xai-org/x-algorithm. Follow QUICKSTART.md in the Phoenix directory to run synthetic training and verify scoring pipelines locally.

Total Cost of Ownership (TCO)

The repository is open source under GNU AGPL v3. Self-hosted compute cost depends on vector index sizing and GPU inference clusters. For standard production applications, self-hosting this infrastructure requires full-time ranking and machine learning operations engineers.

Recommendation

Inspect the repository for systems engineering and multi-stage candidate generation patterns. Use managed recommendation engines or simpler retrieval-augmented pipelines for product MVP feeds unless your business demands petabyte-scale custom feed orchestration.

APPLY ACROSS YOUR WHOLE STACK · $199 USD

Need help architecting recommendation pipelines or vector retrieval feeds?

Reframe ($199) evaluates your recommendation & discovery stack-Rust Home Mixer vs in-memory caching vs vector transformers-auditing latency bottlenecks, cold-start handling, and P99 response times. Diagnosis only.

Fixed $199 fee · 100% vendor-neutral review · 3-day delivery guarantee