Commentary#Agents#MCP#Prompts#Agent memory#Evals

The agent table has flipped: process, memory, and evals beat prompt copying

A WeChat Channels video argues prompt copying is entry-level; value sits in workflows, memory, evals, and tooling. Verified against Anthropic and Cognition.

A wooden chessboard with all pieces laid out, standing in for "the agent table has changed" (Wikimedia Commons, CC0)

On September 19, 2026, the WeChat Channels account 元智社 (Yuanzhishe) posted a short opinion video whose entire caption reads: “The agent table has changed. What’s valuable now is running processes, memory, evaluation, and tooling steadily — copying prompts is becoming an entry-level move.” The video offers no evidence, but its conclusion points the same way Anthropic’s and Cognition’s published engineering guidance does. This piece unpacks that sentence term by term: what holds up, and what deserves a discount.

Prompts are the entry fee

Copying a well-tuned prompt used to be a genuine free lunch, for a simple reason: most tasks then lived inside a single context window and ended after one exchange. Agents don’t. Anthropic put it plainly in Building effective agents (December 19, 2024): “They are typically just LLMs using tools based on environmental feedback in a loop.” Once the loop gets long, the prompt only decides the quality of the first step; every step after that depends on how good the tools are, whether the context still holds, and whether failures get noticed.

Memory and context

Context is a limited resource reshuffled every turn — the part prompts don’t control. In Effective context engineering for AI agents (September 29, 2025), Anthropic named the discipline that extends prompt engineering: “Context engineering is the art and science of curating what will go into the limited context window.” The same post confirms context rot: the more tokens pile into the window, the worse the model recalls any single fact. Memory, then, is budget management rather than an optional extra — persisting key facts as external notes, compacting near-full conversations, giving sub-agents clean windows. That is what “running memory steadily” refers to.

Evals come first

Evaluation is the precondition of “steady”: you define what steady means before you can catch regressions. Anthropic’s Demystifying evals for AI agents (January 9, 2026) states: “Evals make problems and behavioral changes visible before they affect users, and their value compounds over the lifecycle of an agent.” Teams without an eval suite learn about problems from user complaints and cannot tell regressions from noise. Cognition’s Don’t build multi-agents (June 12, 2025) calls context engineering “effectively the #1 job of engineers building AI agents” — and what powers that iteration is a test suite you can reproduce.

Tooling has a standard

The tooling layer already has a de facto standard. The Model Context Protocol shipped from Anthropic on November 25, 2024; OpenAI announced adoption on March 25, 2025, and Google DeepMind followed on April 9, 2025 (per Wikipedia, checked September 28, 2026). Once tool integration standardizes, the moat moves from “I have a secret prompt” to “my tool docs are clear and my failures are diagnosable” — Anthropic says they spent more time optimizing tools than prompts on SWE-bench. VibeKit, a sandbox SDK for coding agents, and Team9, a local-first agent runtime, both covered on this site, are infrastructure in exactly this layer.

The case against over-engineering

The strongest counterargument: for most solo projects, all of this is over-engineering. When the task is narrow and volume is low, prompt tuning still works. Eval suites need maintenance; memory layers like Mem0 and Letta add new failure surfaces; and even Anthropic doesn’t recommend reaching for orchestration frameworks first — the same essay notes, “Consistently, the most successful implementations weren’t using complex frameworks or specialized libraries.” Engineering investment pays back only at scale. The table has changed, but that doesn’t mean everyone should bet now.

One honest caveat: the original video is a single captioned sentence with no evidence chain. Every argument here comes from published engineering guidance; the opinion belongs to 元智社, the verification to this site.

A ten-minute self-test

You don’t need a framework. Run the same input through your agent ten times and count how often the results disagree, and where the drift starts. This ten-minute ritual is an eval suite in embryo, and it will tell you directly whether your bottleneck is the prompt — or the table it sits on.