Back to list
Castform and Neon Outperform GPT-5.6 Sol in Retrieval Accuracy at 100x Lower Cost
Industry NewsAI AgentsOpen SourceDatabase Technology

Castform and Neon Outperform GPT-5.6 Sol in Retrieval Accuracy at 100x Lower Cost

A breakthrough collaboration between Castform and Neon has demonstrated that a 4B open-source model, post-trained with Reinforcement Learning (RL), can match the retrieval accuracy of frontier models like GPT-5.6 Sol. This approach addresses the primary challenges of modern AI agents: providing the right context and enabling intelligent search decisions. By leveraging Neon’s Lakebase Postgres and Castform’s specialized post-training, the solution achieves a 100x cost reduction compared to frontier models. While multi-turn search requests using GPT-5.6 Sol typically cost $0.03 and take over 10 seconds, the Castform-Neon synergy offers a significantly more efficient and scalable alternative for agentic retrieval, marking a shift from traditional RAG pipelines to complex, multi-hop search workflows.

Hacker News

Key Takeaways

  • Performance Parity: A 4B open-source model post-trained with Castform achieves search retrieval accuracy equal to GPT-5.6 Sol.
  • Massive Cost Efficiency: The specialized open-source approach is 100x cheaper than using frontier models for the same tasks.
  • Infrastructure Synergy: The solution combines Neon’s Lakebase Postgres for data context with Castform’s model logic for search decision-making.
  • Evolution of Search: The industry is moving from 2022-era one-shot RAG pipelines to 2025-era multi-hop agentic retrieval loops.
  • Latency Reduction: The new method addresses the >10s latency issues common in multi-turn search requests handled by frontier models.

In-Depth Analysis

The Evolution from RAG to Agentic Retrieval

The landscape of AI-driven data retrieval has undergone a significant transformation over the last few years. In approximately 2022, the industry focused heavily on embedding search, with database providers like Neon integrating extensions such as pgvector to support Retrieval-Augmented Generation (RAG). These early systems were largely "one-shot," where a single query was issued to find relevant context for a Large Language Model (LLM).

By 2025, the paradigm shifted toward agentic search. In this new model, developers create multi-hop workflows where agents decompose complex problems into smaller, manageable tasks. Instead of a single search, the model operates in a loop, planning and searching multiple times to refine results. However, this iterative process introduced a critical bottleneck: every loop iteration required a call to a frontier model, leading to exponential increases in both cost and latency. A typical multi-turn search with a model like GPT-5.6 Sol can cost roughly $0.03 and take more than 10 seconds, making it impractical for high-scale applications.

Bridging the Gap with RL Post-Training

While small open-weights models are inherently 100x cheaper than frontier models, their out-of-the-box capabilities often lag behind closed API models. The collaboration between Castform and Neon proves that this gap can be bridged through Reinforcement Learning (RL) post-training. By post-training a 4B model specifically for retrieval tasks, Castform has enabled it to match the accuracy of much larger, more expensive frontier models.

This strategy focuses on the two core requirements of a "good agent":

  1. Context: The ability to provide tools to find the right data, solved by Neon’s Lakebase Postgres and its new Search extensions.
  2. Model: The ability of the model to decide what to search for, solved by Castform’s post-training logic.

By pointing Castform at Neon, organizations can transform raw database data into usable intelligence without the need for advanced, expensive infrastructure or handcrafted RAG pipelines.

Infrastructure and Efficiency Gains

The primary advantage of the Castform and Neon approach lies in its ability to utilize existing database assets. As Ying Hang Seah, co-founder of Castform, noted, most teams have their best training data sitting idle in databases. The difficulty has always been turning that raw data into something usable for agents to read, search, and mutate at scale.

Neon’s Lakebase Postgres serves as the foundational infrastructure that allows agents to access data efficiently. When combined with a 4B model that has been optimized for decision-making, the system eliminates the prohibitive costs of frontier models. This allows for agentic retrieval that is not only accurate but also fast enough for real-time user requests, overcoming the 10-second latency barrier that currently plagues multi-turn frontier model workflows.

Industry Impact

The success of the Castform and Neon integration signals a major shift in how AI agents will be deployed in the enterprise. By proving that a 4B model can rival GPT-5.6 Sol in specific retrieval tasks, the industry may see a move away from general-purpose frontier models toward specialized, post-trained open-source models. This democratization of high-performance retrieval allows smaller teams to build sophisticated agents that were previously too expensive to operate. Furthermore, the emphasis on "agentic retrieval" over simple RAG suggests that the next generation of AI tools will be defined by their ability to perform complex, multi-step reasoning within a database environment at a fraction of current costs.

Frequently Asked Questions

Question: How does the cost of the 4B model compare to GPT-5.6 Sol?

According to the report, the 4B open-source model post-trained with Castform is 100x cheaper than GPT-5.6 Sol. While a typical multi-turn search with GPT-5.6 Sol costs approximately $0.03, the open-source alternative provides the same accuracy at a fraction of that price.

Question: What is the difference between traditional RAG and agentic retrieval?

Traditional RAG (common around 2022) usually involves a one-shot embedding similarity search to provide context to an LLM. Agentic retrieval (emerging in 2025) involves multi-hop workflows where the model plans and searches multiple times in a loop, decomposing large problems into smaller ones to achieve higher accuracy.

Question: What role does Neon play in this solution?

Neon provides the "Context" component through its Lakebase Postgres and Search extensions. This infrastructure allows agents to find, read, and search the right data efficiently, which is then processed by the Castform-optimized model to make search decisions.

Related News

Industry News

Parallel Cuts Labor Market Research Time and Cost in Half Using OpenAI GPT-6 Astra

According to a release by OpenAI, Parallel has successfully halved both the operational time and overall financial cost required to research and synthesize complex labor-market data by integrating GPT-6 Astra into its agentic workflows. By deploying GPT-6 Astra, Parallel's autonomous agents achieve double the processing efficiency compared to prior models while simultaneously cutting operational expenses by fifty percent. This deployment highlights tangible performance gains in practical agent-driven data analysis and labor research pipelines.

Industry News

OpenAI Outlines Core Priorities and Principles for Rigorous and Independent Third-Party AI Safety Assessments

OpenAI has officially outlined a set of priorities and foundational principles aimed at guiding effective third-party AI safety assessments. As artificial intelligence advances into increasingly capable territory, the organization emphasizes the necessity of independent, rigorous, and secure evaluations targeting frontier models and their corresponding technical safeguards. This initiative highlights the growing recognition across the artificial intelligence sector that internal safety testing alone is insufficient for establishing comprehensive risk mitigation. By formalizing expectations around external assessment methodologies, OpenAI aims to promote transparent verification practices and robust safety validation. The framework addresses the need for external evaluators to thoroughly examine frontier system capabilities and safeguard effectiveness without compromising security, setting a strategic direction for future independent AI auditing standards.

Apple Agrees to $250 Million Siri AI Settlement: Eligible iPhone Owners Can Now Submit Payout Claims
Industry News

Apple Agrees to $250 Million Siri AI Settlement: Eligible iPhone Owners Can Now Submit Payout Claims

Apple has agreed to a $250 million settlement following allegations that the company failed to deliver an advertised AI-upgraded Siri, opening the claims submission process for eligible smartphone purchasers. The resolution allows qualifying United States residents who purchased an iPhone 15 Pro, iPhone 15 Pro Max, or any iPhone 16 model beginning on June 10, 2024, to seek financial compensation through official claims channels. The legal outcome reflects heightened consumer expectations and stricter accountability surrounding marketed artificial intelligence features versus actual product rollouts. This massive financial payout marks an important development for affected consumers and sets a clear precedent for tech companies promoting advanced AI capabilities on flagship hardware.