MLog

All Posts

Filter by keyword, tag, and category.

Tech#模型路由#AI基础设施#LLM#成本优化

Model Routing: Why the Smartest AI Is No Longer the Winner in 2026

In 2026, the defining question for enterprise AI deployment has shifted from 'which model is the smartest?' to 'how do we orchestrate multiple models without going broke?' Model routing can slash inference expenses by 60-80% while maintaining quality. This article examines the evolution from static rules to learned routers, resilience patterns, and Perplexity's hybrid inference as a case study.

Jul 18, 20268 min
Tech#AI基础设施#国产算力#超智融合#高性能计算

Sugon 8000: China's AI Computing Enters the 100,000-Card Era

Sugon launches China's first fully domestic 100,000-card AI supercluster, featuring native supercomputing-intelligence fusion architecture and full-stack self-reliance, pushing AI infrastructure from 10K to 100K scale.

Jul 18, 20264 min
Tech#AI Agent#AI智能体#AI基础设施#开源项目

2026: When AI Agents Started Doing Real Work

WAIC 2026 delivered a clear signal: the AI race has shifted from a parameter arms race to agent-driven engineering. Token calls grew 1,000x in two years, open-source adoption nears a tipping point, and humanoid robots are walking into factories. The 'GPT moment' may have already arrived—just not evenly distributed yet.

Jul 18, 20266 min
Tech#AI开发工具#开源项目#LLM#国产算力

Meituan LongCat-2.0: How a 1.6T Parameter Model Runs on 50,000 Domestic GPUs

Meituan open-sourced LongCat-2.0, a 1.6T-parameter MoE model with 48B active parameters—the first trillion-scale model trained and deployed entirely on Chinese-manufactured GPUs. Here's the architecture, the engineering, and why it matters.

Jul 18, 20264 min
Tech#AI开发工具#AI编程辅助#AI智能体#编程

Stop Prompting, Start Designing: How Loop Engineering is Reshaping AI Programming

In June 2026, Claude Code creator Boris Cherny and Google engineering lead Addy Osmani independently declared a new era: Loop Engineering—where you no longer prompt agents directly, but design systems that prompt themselves. This article breaks down the four paradigm shifts, five core building blocks, and the pitfalls already surfacing in practice.

Jul 18, 20266 min
Tech#AI开发工具#开源项目#模型量化#LLM

From 600GB to 85GB: How Tencent Shrank a 295B Model to Fit on a Single GPU

Eight days after launching Hy3, Tencent's Hunyuan team delivered 1-bit and 4-bit quantized GGUF builds of the 295B flagship model—shrinking it from 598GB to 85.5GiB. Here's what the benchmarks say and why it matters for local AI deployment.

Jul 18, 20265 min
Tech#AI开发工具#AI智能体框架#开源项目#工程化

Superpowers Hits 226K Stars: AI Coding's Paradigm Shift from Code Generation to Workflow Engineering

In July 2026, Superpowers dominates GitHub Trending with 226K stars. It's not a new model—it's a Markdown-based AI coding methodology redefining the developer-AI relationship. The key insight: it's not about making AI smarter, it's about making it follow the rules.

Jul 18, 20266 min
Tech#AI编程辅助#LLM#AI Agent#工程化

When No One Understands the Code: The Trust Crisis in AI-Generated Software

Google generates 75% of its new code with AI. Meta mandates Agent-assisted commits. Yet a Reddit developer watched AI delete 28,745 lines of code without reason, and Moonwell lost $1.78 million to an AI-coded bug. Zuckerberg admits AI Agent progress is behind schedule—the software industry is running naked at full speed.

Jul 18, 20265 min