Back to list
Meituan LongCat Releases General 365: A New Benchmark for AI Reasoning Evaluation
Industry NewsMeituanAI BenchmarkingReasoning Models

Meituan LongCat Releases General 365: A New Benchmark for AI Reasoning Evaluation

The Meituan LongCat team has officially launched General 365, a rigorous new benchmark designed to evaluate the reasoning capabilities of artificial intelligence models. In an initial assessment of 26 mainstream models, the results reveal a significant performance gap in the industry. Google's Gemini 3 Pro, currently regarded as the strongest performer, achieved an accuracy rate of only 62.8%. Notably, the vast majority of the models tested failed to reach the 60% passing threshold, highlighting the intense difficulty of the General 365 evaluation. This release by Meituan sets a new standard for measuring high-level cognitive tasks in AI, suggesting that current large language models still face substantial hurdles in complex reasoning scenarios.

美团技术团队

Key Takeaways

  • New Evaluation Standard: Meituan's LongCat team has introduced General 365, a benchmark specifically focused on reasoning capabilities.
  • Industry Performance Gap: Out of 26 mainstream models tested, the majority failed to achieve a score of 60%.
  • Top Performer: Gemini 3 Pro currently leads the benchmark but only managed an accuracy rate of 62.8%.
  • Rigorous Testing: The benchmark is designed to be a "new yardstick," indicating a higher level of difficulty than previous evaluation methods.

In-Depth Analysis

The Launch of General 365 and the Reasoning Challenge

The Meituan LongCat team has officially released General 365, positioning it as a critical new benchmark for the AI industry. The primary objective of this tool is to provide a "new yardstick" for reasoning evaluation, moving beyond simple task completion to test the underlying logic and cognitive depth of large language models. The introduction of General 365 comes at a time when the industry is seeking more nuanced ways to differentiate between models that can perform basic functions and those that truly possess advanced reasoning skills.

According to the data provided by the LongCat team, the benchmark is intentionally designed to be challenging. By focusing on reasoning, Meituan is targeting one of the most difficult frontiers in AI development. The name "General 365" suggests a comprehensive, perhaps year-round or all-encompassing approach to testing, though the core focus remains strictly on the accuracy of reasoning outputs across a wide variety of scenarios.

Comparative Performance of Mainstream Models

The initial testing phase of General 365 involved 26 of the most prominent AI models currently available in the market. The results of these tests serve as a sobering reality check for the state of AI reasoning. Even the most advanced models struggled to maintain high accuracy levels when subjected to the General 365 criteria.

Gemini 3 Pro, which is identified as the current industry leader in terms of raw performance, reached an accuracy of 62.8%. While this score places it at the top of the list among the 26 models tested, it also highlights how much room for improvement remains. Perhaps more significant is the finding that the "passing line" of 60% was out of reach for the vast majority of models. This failure to meet a basic 60% threshold suggests that many current AI architectures, while proficient in language generation, still lack the robust reasoning frameworks required to navigate the complexities presented by the General 365 benchmark.

Industry Impact

The release of General 365 by Meituan's LongCat team is likely to have a profound impact on how AI models are developed and marketed. By establishing a benchmark where even the strongest models barely exceed a 60% accuracy rate, Meituan is forcing a shift in the industry's focus. Developers may now be incentivized to prioritize reasoning and logical consistency over mere fluency or parameter count.

Furthermore, the fact that a major technology player like Meituan is contributing to the evaluation ecosystem suggests a move toward more transparent and standardized testing. As models continue to evolve, benchmarks like General 365 will be essential for identifying which systems are truly capable of handling complex, real-world problem-solving. This benchmark sets a high bar, serving as both a challenge to current AI leaders and a roadmap for future research and development in the field of artificial intelligence reasoning.

Frequently Asked Questions

Question: What is the General 365 benchmark?

General 365 is a new reasoning evaluation benchmark released by the Meituan LongCat team. It is designed to serve as a rigorous standard for testing the logical and reasoning capabilities of mainstream AI models.

Question: Which model performed the best on the General 365 test?

According to the initial results, Gemini 3 Pro is the top-performing model on the General 365 benchmark, achieving an accuracy rate of 62.8%.

Question: How did most models perform on this new benchmark?

The majority of the 26 mainstream models tested failed to reach the 60% accuracy mark, which is considered the passing line for the General 365 evaluation.

Related News

Industry News

Parallel Cuts Labor Market Research Time and Cost in Half Using OpenAI GPT-6 Astra

According to a release by OpenAI, Parallel has successfully halved both the operational time and overall financial cost required to research and synthesize complex labor-market data by integrating GPT-6 Astra into its agentic workflows. By deploying GPT-6 Astra, Parallel's autonomous agents achieve double the processing efficiency compared to prior models while simultaneously cutting operational expenses by fifty percent. This deployment highlights tangible performance gains in practical agent-driven data analysis and labor research pipelines.

Industry News

OpenAI Outlines Core Priorities and Principles for Rigorous and Independent Third-Party AI Safety Assessments

OpenAI has officially outlined a set of priorities and foundational principles aimed at guiding effective third-party AI safety assessments. As artificial intelligence advances into increasingly capable territory, the organization emphasizes the necessity of independent, rigorous, and secure evaluations targeting frontier models and their corresponding technical safeguards. This initiative highlights the growing recognition across the artificial intelligence sector that internal safety testing alone is insufficient for establishing comprehensive risk mitigation. By formalizing expectations around external assessment methodologies, OpenAI aims to promote transparent verification practices and robust safety validation. The framework addresses the need for external evaluators to thoroughly examine frontier system capabilities and safeguard effectiveness without compromising security, setting a strategic direction for future independent AI auditing standards.

Apple Agrees to $250 Million Siri AI Settlement: Eligible iPhone Owners Can Now Submit Payout Claims
Industry News

Apple Agrees to $250 Million Siri AI Settlement: Eligible iPhone Owners Can Now Submit Payout Claims

Apple has agreed to a $250 million settlement following allegations that the company failed to deliver an advertised AI-upgraded Siri, opening the claims submission process for eligible smartphone purchasers. The resolution allows qualifying United States residents who purchased an iPhone 15 Pro, iPhone 15 Pro Max, or any iPhone 16 model beginning on June 10, 2024, to seek financial compensation through official claims channels. The legal outcome reflects heightened consumer expectations and stricter accountability surrounding marketed artificial intelligence features versus actual product rollouts. This massive financial payout marks an important development for affected consumers and sets a clear precedent for tech companies promoting advanced AI capabilities on flagship hardware.