Back to list
Meituan LongCat Releases General 365: A New Reasoning Benchmark Where Most AI Models Fail to Pass
Industry NewsArtificial IntelligenceBenchmarkingMeituan

Meituan LongCat Releases General 365: A New Reasoning Benchmark Where Most AI Models Fail to Pass

The Meituan LongCat team has officially open-sourced 'General 365,' a rigorous new benchmark designed to evaluate the reasoning capabilities of large language models. In an initial assessment of 26 mainstream AI models, the results highlight a significant gap in current cognitive performance. Even Gemini 3 Pro, identified as the top performer in the test, achieved an accuracy rate of only 62.8%. Furthermore, the vast majority of the models tested were unable to reach the 60% passing threshold. This release by Meituan's technology team provides a new standard for the industry, revealing that complex reasoning remains a substantial challenge for even the most advanced artificial intelligence systems currently available.

美团技术团队

Key Takeaways

  • Meituan's LongCat team has officially released and open-sourced the General 365 reasoning benchmark.
  • Evaluation of 26 mainstream models reveals that most current AI systems struggle with complex reasoning tasks.
  • Gemini 3 Pro emerged as the top performer with an accuracy of 62.8%, yet this remains relatively low for a leading model.
  • The majority of tested models failed to reach a 60% accuracy score, establishing a high difficulty ceiling for the benchmark.

In-Depth Analysis

The Launch of General 365

The Meituan LongCat team has introduced General 365 as a specialized tool for evaluating the reasoning depth of artificial intelligence. By open-sourcing this benchmark, the team provides the global AI community with a new metric to measure progress beyond simple linguistic fluency. The focus of General 365 is specifically on 'General' reasoning, suggesting a broad application across various logical domains. The release comes at a time when the industry is shifting its focus from model size to the quality of logical output and problem-solving efficiency.

Performance Gap in Mainstream Models

The results released alongside the benchmark provide a sobering look at the current state of AI. Out of 26 mainstream models tested, the performance was notably lower than what is typically seen on standard benchmarks. The fact that Gemini 3 Pro, currently regarded as one of the most capable models globally, only secured a 62.8% accuracy rate indicates that General 365 contains highly challenging reasoning problems. This data point serves as a critical indicator that even the 'strongest' models have significant room for improvement when faced with the specific criteria set by the LongCat team.

The 60% Passing Threshold

A striking finding from the LongCat team's report is that the vast majority of the 26 models failed to reach the 60% mark. In many academic and professional contexts, 60% is considered the minimum threshold for a 'passing' grade. The failure of most mainstream models to meet this baseline suggests that current LLM (Large Language Model) architectures may still lack the robust logical frameworks required for consistent reasoning. This gap between current performance and the 'passing' line highlights the rigorous nature of General 365 as a evaluative standard.

Industry Impact

The introduction of General 365 by Meituan is significant for the AI industry as it establishes a more demanding yardstick for reasoning. By making the benchmark open-source, Meituan allows other developers to stress-test their models against the same 26-model baseline. This could lead to a shift in development priorities, moving away from general knowledge retrieval and toward the enhancement of internal logic and multi-step reasoning. As models strive to exceed the 62.8% mark set by Gemini 3 Pro, General 365 will likely become a key reference point for future iterations of large language models.

Frequently Asked Questions

Question: What is the significance of the 62.8% score achieved by Gemini 3 Pro?

Within the context of the General 365 benchmark, 62.8% represents the highest accuracy among 26 mainstream models. While it leads the field, the score suggests that even top-tier AI models face difficulty with the reasoning tasks included in this specific evaluation.

Question: Why did most models fail to reach the 60% mark on General 365?

The failure of the majority of models to reach 60% indicates that the General 365 benchmark is designed with a high level of difficulty that targets the weaknesses in current AI reasoning capabilities, setting a new and more difficult standard for the industry.

Question: Is General 365 available for public use?

Yes, the Meituan LongCat team has officially open-sourced General 365, allowing the broader technology community to use it for evaluating and improving AI model reasoning.

Related News

Industry News

Parallel Cuts Labor Market Research Time and Cost in Half Using OpenAI GPT-6 Astra

According to a release by OpenAI, Parallel has successfully halved both the operational time and overall financial cost required to research and synthesize complex labor-market data by integrating GPT-6 Astra into its agentic workflows. By deploying GPT-6 Astra, Parallel's autonomous agents achieve double the processing efficiency compared to prior models while simultaneously cutting operational expenses by fifty percent. This deployment highlights tangible performance gains in practical agent-driven data analysis and labor research pipelines.

Industry News

OpenAI Outlines Core Priorities and Principles for Rigorous and Independent Third-Party AI Safety Assessments

OpenAI has officially outlined a set of priorities and foundational principles aimed at guiding effective third-party AI safety assessments. As artificial intelligence advances into increasingly capable territory, the organization emphasizes the necessity of independent, rigorous, and secure evaluations targeting frontier models and their corresponding technical safeguards. This initiative highlights the growing recognition across the artificial intelligence sector that internal safety testing alone is insufficient for establishing comprehensive risk mitigation. By formalizing expectations around external assessment methodologies, OpenAI aims to promote transparent verification practices and robust safety validation. The framework addresses the need for external evaluators to thoroughly examine frontier system capabilities and safeguard effectiveness without compromising security, setting a strategic direction for future independent AI auditing standards.

Apple Agrees to $250 Million Siri AI Settlement: Eligible iPhone Owners Can Now Submit Payout Claims
Industry News

Apple Agrees to $250 Million Siri AI Settlement: Eligible iPhone Owners Can Now Submit Payout Claims

Apple has agreed to a $250 million settlement following allegations that the company failed to deliver an advertised AI-upgraded Siri, opening the claims submission process for eligible smartphone purchasers. The resolution allows qualifying United States residents who purchased an iPhone 15 Pro, iPhone 15 Pro Max, or any iPhone 16 model beginning on June 10, 2024, to seek financial compensation through official claims channels. The legal outcome reflects heightened consumer expectations and stricter accountability surrounding marketed artificial intelligence features versus actual product rollouts. This massive financial payout marks an important development for affected consumers and sets a clear precedent for tech companies promoting advanced AI capabilities on flagship hardware.