Back to List
Meituan LongCat Releases General 365 Reasoning Benchmark as Leading AI Models Struggle to Pass
Industry NewsMeituanArtificial IntelligenceBenchmarking

Meituan LongCat Releases General 365 Reasoning Benchmark as Leading AI Models Struggle to Pass

The Meituan LongCat team has officially launched General 365, a rigorous new benchmark designed to evaluate the reasoning capabilities of large language models (LLMs). In a comprehensive test involving 26 mainstream AI models, the results revealed a significant performance gap in the industry. Even the high-performing Gemini 3 Pro, currently regarded as one of the most capable models available, achieved an accuracy rate of only 62.8%. Furthermore, the evaluation demonstrated that the vast majority of tested models were unable to reach the 60% accuracy threshold, which is traditionally considered a passing grade. This release by Meituan's technology team establishes a challenging new standard for AI reasoning, highlighting that current frontier models still face substantial hurdles in mastering complex logical tasks.

美团技术团队

Key Takeaways

  • Launch of General 365: Meituan's LongCat team has introduced a new evaluation standard specifically focused on reasoning capabilities.
  • Widespread Performance Gap: Out of 26 mainstream models tested, most failed to achieve a 60% accuracy rate.
  • Gemini 3 Pro Results: The industry-leading Gemini 3 Pro secured the top spot but only managed a 62.8% accuracy score.
  • New Industry Benchmark: General 365 is positioned as a high-bar metric that exposes the limitations of current large language models in complex reasoning scenarios.

In-Depth Analysis

The Challenge of General 365

The introduction of General 365 by the Meituan LongCat team marks a pivotal shift in how artificial intelligence reasoning is measured. By testing 26 of the most prominent models in the current market, the benchmark provides a sobering look at the state of AI development. The fact that a majority of these models could not surpass the 60% mark suggests that General 365 is designed to test depth and logical consistency rather than simple pattern matching. This high level of difficulty serves to differentiate truly capable reasoning engines from those that rely on surface-level heuristics.

Benchmarking the Frontier: Gemini 3 Pro

One of the most significant findings from the Meituan report is the performance of Gemini 3 Pro. Despite its reputation as a leading model in the global AI landscape, its accuracy on the General 365 benchmark was limited to 62.8%. While this score placed it at the top of the 26 models tested, the narrow margin by which it passed the 60% threshold indicates that even the most advanced systems have considerable room for improvement. This data point underscores the rigor of the General 365 evaluation framework and suggests that the "reasoning" capabilities of modern LLMs are still in a relatively early stage of evolution when subjected to such stringent criteria.

The 60% Threshold and Model Failure Rates

The report highlights a concerning trend: the "passing line" of 60% remains out of reach for the bulk of the AI industry. With 26 mainstream models under review, the failure of the majority to reach this basic benchmark suggests a systemic challenge in current model architectures or training methodologies. Meituan's findings imply that while models are becoming more conversational and versatile, their ability to navigate the specific logical complexities demanded by General 365 remains a significant bottleneck. This creates a clear roadmap for future research, emphasizing the need for more robust reasoning frameworks.

Industry Impact

The release of General 365 by Meituan is likely to have a profound impact on the AI research community. By establishing a benchmark where even the strongest models struggle, Meituan has effectively raised the ceiling for what is considered "advanced" reasoning. This move encourages developers to move beyond traditional benchmarks that may have become saturated or prone to data contamination.

Furthermore, the transparency of these results—showing that most mainstream models fall below a 60% accuracy rate—provides a realistic baseline for enterprise expectations. As companies look to integrate AI into complex decision-making processes, benchmarks like General 365 offer a more accurate reflection of a model's reliability in high-stakes reasoning tasks. This will likely drive a new wave of optimization focused specifically on the logical gaps identified by the LongCat team.

Frequently Asked Questions

Question: What is the primary focus of the General 365 benchmark?

General 365 is an open-source benchmark released by the Meituan LongCat team specifically designed to evaluate and set a new standard for the reasoning capabilities of large language models.

Question: How did the top-performing models fare on this benchmark?

According to the test results of 26 mainstream models, Gemini 3 Pro was the top performer with an accuracy of 62.8%. However, the majority of the other models tested failed to reach the 60% accuracy mark.

Question: Why is the 60% accuracy mark significant in this report?

The 60% mark is described as the "passing line." The fact that most mainstream models failed to reach this level highlights the extreme difficulty of the General 365 benchmark and the current limitations of AI reasoning.

Related News

AI in Finance: The Next Major Industry Vertical Following the Success of Coding
Industry News

AI in Finance: The Next Major Industry Vertical Following the Success of Coding

Artificial intelligence is rapidly expanding its footprint within the financial services sector, positioning it as the next primary vertical for AI integration following its transformative impact on software coding. This shift highlights a strategic move toward industry-specific AI applications. Alongside this trend, the opening of AIE NYC marks a significant milestone in establishing dedicated hubs for AI development. This analysis explores the transition of AI from programming tools to financial systems and the implications of localized AI initiatives like AIE NYC in driving the next wave of technological adoption in the finance industry.

Mark Zuckerberg Forecasts Billions of Personal AI Agents Within Five Years Amid Massive Meta Infrastructure Investment
Industry News

Mark Zuckerberg Forecasts Billions of Personal AI Agents Within Five Years Amid Massive Meta Infrastructure Investment

Meta CEO Mark Zuckerberg has issued a bold prediction stating that billions of people will utilize personal AI agents within the next five years. This forecast comes at a time when Meta is directing billions of dollars into AI infrastructure and the development of specialized agents. Zuckerberg's primary objective is to demonstrate to investors that these substantial capital expenditures will result in a significant long-term payoff. The vision centers on a future where AI agents are a ubiquitous part of the human experience, supported by a massive technological foundation currently being built by Meta. The five-year timeline sets a specific horizon for the industry to transition from experimental AI tools to widespread, personal agentic systems used on a global scale.

Microsoft Reports $3.2 Billion Gain from Anthropic Investment Amid Mixed OpenAI Financial Results
Industry News

Microsoft Reports $3.2 Billion Gain from Anthropic Investment Amid Mixed OpenAI Financial Results

Microsoft's fiscal year 2026 fourth-quarter earnings report has revealed a significant $3.2 billion gain from its investment in Anthropic. While the company celebrated overall strong financial performance, the report characterized its investment in OpenAI as a "mixed bag." This disclosure, tucked into the year-end results ending June 30, provides a rare financial comparison between Microsoft's stakes in the two primary competing AI laboratories. The contrast highlights the varying financial trajectories of the industry's leading AI developers and Microsoft's strategic positioning as a major backer of both rivals. The findings suggest a complex financial dynamic as Microsoft navigates its partnerships with the most prominent entities in the artificial intelligence sector.