Perplexity has introduced Photon, a new Rust-based retrieval and ranking system designed to make its search infrastructure faster and more efficient, alongside a new Fast Search option for the Perplexity Search API. The company says Fast Search can return 95% of results within 230 milliseconds while cutting the cost of agent tasks by 68% compared with its default preset.
The infrastructure change is aimed primarily at developers building AI agents and applications that make frequent web-search calls. Perplexity says Photon now handles retrieval and ranking across its search systems, replacing the open-source engine it previously used.
The announcement expands on Perplexity’s longer-term effort to build search infrastructure specifically for AI workloads. Its technical research describes a search architecture combining hybrid retrieval, multi-stage ranking, distributed indexing and AI-assisted content understanding across an index containing more than 200 billion unique URLs.
Quick summary
- Perplexity Photon is a new Rust-based retrieval and ranking system developed by Perplexity.
- It powers Fast Search in the Perplexity Search API.
- Perplexity reports significantly lower search latency and infrastructure requirements.
- Photon replaces the company’s previous retrieval engine across its search infrastructure.
- Its architecture focuses on compact indexes, efficient data reads and faster ranking.
- Fast Search is designed particularly for agentic and high-volume search workflows.
- Perplexity reports some trade-offs in long-tail relevance and answer availability.
What Perplexity Fast Search Changes?
Fast Search is positioned as a lower-compute search configuration for applications where response speed and search cost are particularly important.
According to Perplexity’s reported internal testing, Fast Search returned 95% of results within 230 milliseconds. Its median latency was approximately 160 milliseconds, which the company says was the lowest among the search APIs included in its comparison.
Perplexity also reports a 68% reduction in cost per task across six agent benchmarks compared with its default Search API preset, while maintaining comparable overall quality.
The trade-off is that Fast Search uses less compute during ranking. Perplexity’s internal tests found a 0.24-point reduction in relevance and roughly a three-percentage-point reduction in answer availability on its long-tail and broad-coverage query evaluations.
That makes Fast Search less of a universal replacement for higher-quality search configurations and more of a performance-oriented option for routine retrieval.
Perplexity’s developer platform describes its Search API as a low-latency hybrid search system combining semantic methods, LLM ranking and human feedback, with ranked results, filtering and content extraction available to developers.
Photon Replaces Perplexity’s Previous Retrieval Engine
Photon is the infrastructure layer behind the performance changes.
Perplexity describes it as a Rust-based retrieval and ranking service developed by a small engineering team with extensive use of AI agents during development. The system now handles retrieval and ranking throughout Perplexity’s search infrastructure.
One of the central design decisions is how Photon stores and reads its search index.
Rather than requiring every query to process large amounts of stored information, Photon uses compact index formats so each query reads and decodes only the information it needs. When data must be retrieved from disk, the system batches those reads and performs them asynchronously.
This architecture is intended to prevent storage access from becoming a bottleneck in the search path.
The approach fits with Perplexity’s previously documented search architecture, which separates broad candidate retrieval from progressively more sophisticated ranking stages. Its technical overview describes a pipeline that combines lexical and semantic retrieval before applying multiple ranking stages, including embedding-based scorers and more computationally expensive cross-encoder rerankers.
Photon Also Changes How Search Infrastructure Is Updated
Photon is designed to allow search indexes to be updated without taking the entire serving system offline.
Perplexity says new index versions are built separately from the machines handling live queries. When an updated version is ready, the system updates one serving group at a time.
The new serving group is then warmed using real search queries and checked for readiness before traffic is routed to it. Other serving groups continue handling requests during the process.
This approach reduces the need for a single large cutover and allows Perplexity to progressively introduce new index versions while maintaining live search traffic.
That matters at Perplexity’s scale because its search infrastructure is continuously dealing with new and changing web content. The company’s earlier technical work describes an indexing system processing tens of thousands of indexing operations per second and using machine-learning models to decide which pages should be indexed or refreshed.
Perplexity Reports Major Latency Improvements Inside Photon
Perplexity says the internal p99 response time for its retrieval and ranking system fell from approximately 800 milliseconds to about 65 milliseconds after the Photon work.
The company also reports that Photon uses around 20% fewer serving machines while storing approximately 2.5 times more data per document.
These figures describe Perplexity’s internal infrastructure rather than a standardized third-party benchmark, so they should be understood as company-reported engineering results.
The distinction is important because search latency can vary substantially depending on query type, geographic region, workload, network conditions and the amount of downstream processing required. Perplexity’s earlier Search API evaluation, for example, measured API latency from AWS us-east-1 and reported a 358-millisecond median for its broader Search API configuration.
Photon’s reported internal numbers therefore should not be interpreted as a direct replacement for independent end-to-end API measurements.
Why Faster Search Matters for AI Agents?
The main significance of Fast Search is not simply that individual web searches become faster.
AI agents can perform many searches during a single task. An agent researching a company, comparing products, investigating a technical question or gathering information for a report may issue numerous retrieval requests before producing a final response.
In those workflows, search latency and cost can accumulate across multiple tool calls.
A lower-latency retrieval system can reduce the time spent waiting between agent steps, while a cheaper search configuration can make high-volume retrieval more practical.
This is also why Perplexity has increasingly treated search infrastructure as a core part of its AI platform rather than simply as a conventional web-search backend. Its Search API is specifically designed to provide ranked web results and extracted content to developers building applications and AI systems.
Fast Search Does Not Eliminate the Quality Trade-Off
Perplexity’s own measurements show that Fast Search makes a deliberate quality trade-off.
The company reports slightly lower relevance and answer availability on its internal long-tail and broad-coverage evaluations. For routine agent tasks, that difference may be acceptable when speed and cost are more important than squeezing out the highest possible retrieval quality.
For complex research tasks, however, developers may still prefer configurations that allocate more compute to retrieval and ranking.
This distinction is important because the announcement is not simply about making every Perplexity search faster. It represents a move toward different search performance tiers for different AI workloads.
Photon provides the underlying infrastructure, while Fast Search exposes one way of using that infrastructure with a stronger emphasis on latency and efficiency.
What Photon Means for the Perplexity Search API?
The Photon rollout is another step in Perplexity’s effort to build search infrastructure around the requirements of AI agents.
The company’s earlier Search API architecture already emphasized hybrid retrieval, fine-grained document segmentation, multi-stage ranking and continuously refreshed indexing. Photon extends that infrastructure focus toward serving efficiency, storage performance and faster ranking.
For developers, the practical change is the availability of a faster search option for workloads where milliseconds and per-task search costs matter. For Perplexity, the larger infrastructure change is the replacement of its previous retrieval and ranking engine with a system designed specifically around its current search scale.
The detailed performance figures remain Perplexity’s own measurements, but the broader development is consistent with the company’s publicly documented strategy of building its own search stack for real-time AI retrieval. The Perplexity Search API remains available to developers as part of its API platform.
Also Read-
Perplexity AI Guide: Features, Search & How It Works?
Perplexity Computer Effort Controls for AI Tasks
Source
Perplexity – Photon announcement
Perplexity – Architecting and Evaluating an AI-First Search API
Perplexity API Platform – Search API


