Phancy Leads China AI Platform Review with Full Marks in Four Areas

IDC rates Phancy first among Chinese AI resource managers, noting full scores in four key dimensions for heterogeneous chip support.
Key points
- Phancy ranked first in IDC's China AI platform assessment, scoring full marks in four of six dimensions.
- The platform manages chips from over ten vendors, supporting standard frameworks without code changes.
- IDC notes AI infrastructure is evolving to support agentic workloads, requiring unified management and cost control.
Phancy has been ranked first overall in a recent assessment of AI computing resource management platforms in China. The evaluation, conducted by IDC and reported by DagangNews, highlighted the company’s technological maturity in managing diverse hardware and large-scale clusters.
The report notes that Phancy achieved full scores in four out of six evaluated dimensions. This performance indicates a strong capability in heterogeneous accelerator compatibility and hybrid cloud delivery, outperforming the industry benchmark in these specific areas.
Shifting focus from model creation to deployment
According to the IDC report, the AI industry is moving beyond the initial phase of model development. The focus is now shifting toward large-scale deployment, where AI agents are reshaping cloud infrastructure requirements for enterprises.
For many businesses, the challenge is no longer access to large language models. Instead, the priority is translating these capabilities into sustained productivity while maintaining operational stability and controlling costs.
Managing mixed hardware without code changes
Phancy’s platform supports chips from over ten different vendors, allowing for unified management of training and inference tasks. A key feature is code-compatible multi-chip virtualization, which lets standard frameworks like PyTorch run without application code modifications.
This approach enables the partitioning of a single accelerator across multiple workloads. It helps maintain stability when different tasks run concurrently, reducing the need for specialized software for each hardware type.
Balancing flexibility with cost control
The platform uses three-tier resource pools to schedule workloads based on priority and real-time load. This allows large-scale training and high-concurrency inference to share the same infrastructure efficiently.
Cost tracking is supported through dual-track metering, measuring both accelerator hours and token consumption. While this provides detailed visibility, the trade-off is the need for sophisticated configuration to manage quotas and prevent usage spikes.






