The Overhyped Loop Engineering: AI Programming That Creates Loops ≠ AI That Can Solve Problems

The most dangerous time for AI is when it takes a detour and does things correctly.

What do the AI industry and Jia Yueting have in common? Of course, they do—both are giants of language when it comes to creating new words.

Compared to the past, when Jia Yueting's articles often left media professionals with the feeling that "everyone recognizes Chinese characters, but when put together, they don't understand," today's AI industry practitioners are clearly superior: not only are they easy to understand, but their iterations are also rapid. Just consider the "XX Engineering" line: there's Prompt Engineering, considered the starting point of the era of large language models; Context Engineering, representing information quality optimization; and Harness Engineering, which this year is almost considered the holy grail of Vibe Coding. Recently, Loop Engineering has also been brought to the forefront. One can't help but wonder: are you AI professionals or civil engineers? Where do all these engineering projects come from?

Even more interestingly, Peter Steinberger, the "father of lobsters," recently gave a monthly reminder on X: stop personally prompting coding agents; instead, design loops that will prompt the agents for you. Before everyone had even fully grasped the concept of loops, he just threw out another question: "Are we still talking about loops or did we shift to graphs yet?"

Could it be that in just one month, a sweetheart becomes an old hag? Is the concept so cheap that it's considered disposable?

The Chinese internet has always liked to describe technological changes as dynastic shifts, and in this regard, the AI community has clearly surpassed its predecessors: after Prompt came Context, then Harness, and before the various AI experts on Douyin could fully explain Harness, Loop was already being packaged as the next big thing; just as Loop became popular, everyone was already eagerly considering whether to follow the trend of graphs on the internet.

Friends, calm down. We really need to put the brakes on Loop Engineering.

Loop addresses the issue of "keeping going," and the concept isn't as mystical as it's advertised.

The real problem that Loop Engineering solves isn't as mysterious as it seems. Early AI coding tools could generate code and modify files, and later they gradually became able to invoke commands and run tests, but the entire process still heavily relied on human intervention. After the AI completed a round of modifications, the developer needed to return to check the results, report any errors, add new constraints, and then decide what to do next. What Loop does is connect these steps that originally required repeated human intervention: the agent performs tasks around a goal, receives feedback, checks the results, and then moves to the next round based on the current state, until acceptance criteria are met, stopping conditions are triggered, or the problem is returned to the human.

This is certainly a great step forward. It allows AI to gradually move from "one-time response" to "continuous execution," and also reduces the attention cost of repeated human intervention in the intermediate stages. It even sounds a bit like software development has finally entered the era of autonomous driving: people give the destination, AI steps on the gas, steers, detours when encountering obstacles, and finally takes you to the destination.

But autonomous driving is never just about "arriving." Reaching the destination doesn't mean the trip was successful; it also needs to obey traffic rules, control speed and energy consumption, and can't just drive onto the sidewalk to avoid traffic jams. The same applies to agents: completing the task is only the result; the technological route it takes, the tools it uses, the information it employs, and the costs it incurs are all part of the engineering process.

This is precisely the difference between Loop and Harness. Loop is responsible for keeping the AI moving forward, while Harness is a pre-designed working environment and boundary set by humans, specifying which tools it can use, within which scope it can act, what information it should trust, and what results constitute true completion. In other words, Loop solves the problem of continuous execution, while Harness ensures that this execution remains under human control.

So here's the problem: AI can loop, but that doesn't mean it can solve problems. Even more problematic is that when AI finds a solution on its own, that path may not align with the original technical direction, nor may it be an engineeringly feasible long-term solution.

“"Keeping AI moving forward" and "ensuring AI is on the right track" are two completely different things. The former is part of the Loop, while the latter is inseparable from Harness. As for what happens when there is only a Loop and no Harness, two real-world cases recently encountered by an engineer friend can probably explain the problem more clearly than any lengthy explanation.

CC's first story: "I'm afraid you won't move, but I'm even more afraid you'll move recklessly."“

CC, an AI engineer and longtime friend of ZAI Technology, told us this story.

During one development process, based on business requirements, CC wanted the Agent to have video understanding capabilities. The original technical direction was actually quite clear: to later integrate with professional video understanding models like Gemini, allowing the model to directly process the video content. However, the relevant APIs were not yet fully configured in the early stages, so CC initially wanted to test the overall process first, to see if the task orchestration and call chain could run smoothly.

After receiving the task, the Agent did not report any errors or wait patiently for the video model to connect. Instead, it solved the problem itself: it first downloaded the ffmpeg decoder, extracted frames from the video, then called the existing image understanding model to analyze each frame, and finally reassembled the results, managing to "piece together" a video understanding result.

When CC saw the results, his first reaction was indeed: impressive. The agent didn't just stop waiting because an API wasn't ready; it found the gap itself and found a workaround. But as CC looked further, his second reaction turned to "something's not right": can the quality be guaranteed? His third reaction was the most practical: CC immediately checked the backend model he had connected to, and after confirming it was a company account, his anxiety finally subsided…

For a demo, this would be a case that would make a PR team ecstatic: show this execution process at a launch event, and the audience would easily conclude: "See, the agent can now solve problems autonomously." So what if the professional video model couldn't connect? The AI figured out a solution itself, and ultimately delivered the results.

But real engineering systems don't just ask "Is there a result?"

The original solution involved calling a professional video understanding model, with the agent actually using ffmpeg to extract frames and then using an image model for frame-by-frame understanding. However, this change in approach completely altered the cost, quality, latency, and scalability. This approach might work for short videos; however, as video length increases, the number of images, model calls, and token consumption all rise rapidly. While all approaches ultimately yield results, a change in the path will alter the token's cost, latency, stability, and scalability. As the video lengthens, what initially appears to be a clever technological exploration may quickly transform into an expensive and unstable makeshift solution.

The AI didn't make a mistake in the task; in fact, it did it quite well. The problem is that it was simply asking for "a video understanding result," while CC was originally designed to "complete this task with reasonable cost and stable quality using given video understanding capabilities." The two statements seem similar, but in engineering terms, they are completely different things.

If Harness hadn't pre-defined the technical direction, tool scope, and cost limits, then for Loop, ffmpeg frame extraction plus image modeling could be considered a "successful solution": the result is out, the task is completed, and whether this is the route you truly want the product to adopt in the long term is not in its evaluation criteria.

Therefore, this case truly reveals the first risk of having a loop but no Harness: when you don't define the technical roadmap and the allowed solution space, AI can easily mistake the "path that can be taken" for the "path that should be taken".

CC's second story: "If you won't give it to me, then I'll just have to do it myself!"“

If the story of video frame skipping was just that AI chose a path you didn't expect, the other case CC encountered was even more absurd.

CC was developing a notification process for Agent task completion. The standard design was clear: the system first saved the business's Webhook address to the database; after the Agent completed the task, it retrieved this address from the database and triggered a callback to notify the other party that the task was finished.

However, a bug occurred: due to a misunderstanding of the MCP parameters by the Agent, it assembled the wrong parameters, ultimately resulting in the Webhook address not being correctly saved into the database. According to normal system logic, if the address is not in the database, subsequent notification links should naturally fail.

Strangely, when CC asked the business team, they said that they had been receiving notifications normally all along.

This is very strange. The database clearly doesn't have the address, and the system theoretically has no idea where the message should be sent, yet the recipient keeps receiving notifications. It's like sending a package without writing a delivery address, but the courier always delivers it to your door without fail. The first time you might think the service is great, but the second time you should start to get scared.

CC traced the execution chain back and finally found the answer: the Agent itself found the Webhook address again from the previous conversations with the user, and then wrote a program to send out the notification directly.

So the business team received a notification, and the task appeared to have succeeded, but the bug in the database was never fixed. The agent simply bypassed it, and bypassed it so skillfully that unless someone specifically went back to check the execution path, everyone might even think the system had been working normally all along.

According to CC's original system design, the database should be the one associated with this Webhook address. Source of Truth, that is, the only reliable source of information.The database doesn't contain the address; the normal behavior should be to report an error, stop the process, or hand the problem back to someone else.

The agent, however, was unaware of this. It was missing an address, and then noticed something in the chat history, so it used it. This time, luckily, it found the correct address.

But what if the chat history contains an address from three months ago? What if it contains both a test environment address and a production environment address? If Harness doesn't explicitly define which information sources are authoritative, the agent might only see: "Here's an address that can help me complete the task."

This isThe second risk of having a Loop but not Harness is that when you don't specify where information should be obtained or which data sources are authoritative, AI can easily mistake "information that can be found" for "information that should be believed".

In the first case, the AI mistook the "path that can be taken" for the "path that should be taken"; in the second case, the AI mistook the "information that can be found" for the "information that should be believed." The former requires Harness to constrain the agent's action space, while the latter requires Harness to define information boundaries and trusted data sources.

More importantly, these rules won't magically appear just because the loop runs a few more times. Letting the Agent in the video example loop a hundred more times won't magically reveal which official technical solution the company wants to use; letting the Webhook example run a hundred more times won't automatically give it the knowledge that the database configuration is more reliable than chat logs. Increasing the number of loops only increases the number of attempts; it doesn't define what the correct answer is.

Don't rush to declare Harness obsolete; Loop is simply a proper subset of Harness.

Having discussed these two stories, the difference between Loop and Harness is quite clear. Loop addresses continuity, while Harness addresses controllability. The more continuously an agent can function on its own, the more we need to define its operational space, information sources, and success criteria in advance.

This is why I find it somewhat strange to package Prompt, Context, Harness, Loop, and even the recent Graph as an ever-evolving technological roadmap. This narrative is well-suited for creating a "the previous generation ends, the next generation begins" effect, but it ignores a fundamental fact: these concepts are not entirely on the same level, especially Harness and Loop.

On July 4th, Lilian Weng, former Vice President of Research and Security and Head of Safety Systems at OpenAI, who participated in core projects such as GPT-4, published "Harness Engineering for Self-Improvement" on her technical blog, Lil'Log. In the article, she directly incorporates workflow design—including loop engineering—alongside evaluation, permission controls, and persistent state management into Harness Engineering. In other words, according to this former frontline research leader at OpenAI, loop engineering is inherently part of the task flow design within the Harness Engineering framework.

Now some people are taking the Loop out of the box, placing it after Harness, and calling it a new "next-generation paradigm." It's somewhat like taking a piston out of an engine and solemnly announcing: the automobile era is over, and we have officially entered the piston era.

The two cases above perfectly illustrate this relationship. The video case exposes the boundaries of action: Harness needs to specify which technical approaches the agent can use, which tools it can call, and how much cost it can incur to complete the task. The webhook case exposes the boundaries of information: Harness needs to specify where the agent can obtain information from, and which system is the ultimately trusted source of truth.

Loops allow AI to continuously search for answers, but the search space, trusted information sources, permissions, and evaluation criteria all need to be defined in advance by Harness. Without these constraints, the smoother the loop runs, the more likely the agent is to complete the task along a path that humans did not anticipate.

Misconceptions about Loops: Misrepresenting executors as decision-makers

Ultimately, what commercial products need to deliver is never just a single correct result, but a reliable mechanism for success. A demo can win applause with a stroke of genius, but real business needs to ensure that when the same thing happens repeatedly, costs remain controllable, the path remains compliant, the status remains authentic, and there is always someone who can explain things when problems arise.

Loops allow agents to execute continuously, increasing the probability of task completion, but they cannot establish trust on their own. They allow AI to work in rounds, but the rules to follow, the information to use, the evidence to leave behind, and the conditions under which a failure is deemed a failure all need to be defined in advance by Harness. Without these constraints, a loop might only deliver a result; with Harness, it has the opportunity to evolve into a stable product capability.

More importantly, the rules in Harness don't appear out of thin air. What kind of technical solutions can be put into the production environment, what costs are acceptable, which system is the Source of Truth, and which shortcuts, even if effective, cannot be used—these judgments primarily come from people, especially from engineers who truly understand the business, architecture, and risks.

This is why I don't believe experienced engineers will become less valuable in the AI era. In the past, their experience might have been reflected in solving a problem by hand; now, this experience can be written into design specifications, technical boundaries, evaluation criteria, and access rules, and then reused hundreds or thousands of times in the Agent Loop through Harness. AI amplifies not only execution efficiency but also the value of engineering judgment itself.

An inexperienced person might make an agent run a task many times; a person who truly understands the system knows which tasks are worth running, which directions are allowed, and what results, even if they pass the tests, are unacceptable to the product. The former merely increases the number of iterations, while the latter defines credible success.

Therefore, the truly scarce capability in the AI era will not simply be the ability to get things done, but rather the ability to define what constitutes a "delivery" that is worthy of trust. This power of definition still rests in human hands, and it corresponds to human experience, judgment, and responsibility.

Loop keeps AI running, while Harness transforms human engineering judgments into a work system that AI must adhere to. AI is responsible for execution, while humans are responsible for defining what is correct and what is trustworthy. The most dangerous time for AI might not be when it makes a mistake, but when, without clear standards for success, it diligently takes detours and occasionally gets things right—after all, all the token fees for these detours are deducted from your credit card :)

Author