ai technology

Generation Is Cheap, Judgment Is Expensive

Junyoung Park · 2026-08-03 · 20 min

The future did not arrive in the form I had imagined. I had pictured AI suddenly taking people's jobs and overturning the world, but the change I actually encountered was far quieter and more ordinary. AI first began to blur, one by one, the boundaries that had defined my work. Reading papers, writing code, training models, interpreting experimental results, preparing reports, and persuading other people—things once regarded as separate abilities—started to converge within a single interface. I studied AI, and that led to a career researching and developing it. In some respects, I am among the most direct beneficiaries of this era. A technology I had studied with interest for years began receiving broad social attention, and the ability to work with it became a profession. Yet, ironically, the advance of AI gave me a job while continually making me question what that job means.

Training a model once had a fairly high barrier to entry. You had to read papers, understand architectures, prepare data, and solve the many problems that arose during implementation. Even knowing what to suspect when training failed required experience. To someone else, the problem might have looked like nothing more than a line of code that would not run; inside it, however, were entangled the distribution of the data, the characteristics of optimization, the model architecture, and hardware constraints. Things are very different now. LLMs, initially described as autoregressive models that predict the next token, can now read papers, write code, fix errors, plan experiments, and even help train other models. One well-trained object is lowering the barrier to training another. The boundaries around “training AI” and “designing good experiments,” once treated almost as the exclusive domain of researchers, are rapidly collapsing as well.

This is, of course, a positive change too. Technology has become accessible to more people, and fewer people with good ideas are unable even to begin because they lack implementation skills. The democratization of technology is welcome in itself. If a researcher's expertise rested only on being able to use tools unavailable to everyone else, perhaps that expertise was never likely to last. Even so, some discomfort remains. As the technical barrier has fallen, have people come to understand the technology better? Often, the opposite seems true. Before understanding what works and how, people ask how it can be used and described so that it appears more impressive. Adopting something quickly has become more important than understanding it precisely, and the ability to try a new tool before others sometimes earns more visible credit than the willingness to explore it deeply.

The target of persuasion is not always human, either. Writing is shaped to perform well in search engines and recommendation algorithms, while careers and accomplishments are packaged for easy consumption on social networks such as LinkedIn. Papers are no longer addressed only to human readers. AI systems assist with reviews, and AI is deeply involved in writing the papers themselves. One day, researchers may write not to persuade their peers, but to persuade a model trained on historical acceptance decisions and review patterns. Research could then become less a process of discovering new facts than an optimization problem for passing through an existing evaluation distribution: write the abstract in a structure the model will favor, emphasize contributions it is likely to deem important, and arrange the logic and sentences in ways it can readily accept. Search-engine optimization has already changed the form of writing; there is no guarantee that optimization for AI evaluation will not change the form of research.

We question other people's claims relatively easily. We consider their interests, their biases, and why they reached a particular judgment. Strangely, however, we often feel less resistance toward answers produced by an LLM—especially in specialist fields we do not know well. If the sentences are fluid and the logic neatly arranged, we tend to accept them before checking what omissions and biases may be hidden inside. LLMs can seem to compress and present the accumulated data and common thought of humanity, but that generality is never neutral. An average always includes some things and erases others. Heavily documented viewpoints remain stronger; experiences that were not adequately recorded grow weaker. Judgments that are already mainstream are reproduced more naturally, while unfamiliar challenges may be treated as abnormal or unpersuasive. Yet we are gradually losing both the time and the critical capacity to confront the biases concealed within that average. We review and approve AI-generated results in the name of accountability and oversight, but it is difficult to know how deeply we are actually judging them. If the person who presses the final approval button lacks the ability to review the result, enough time to do so, and the authority to reject it when something is wrong, can that really be called human oversight?

The presence of a human in the loop does not guarantee that a human is thinking. AI produces a judgment, a person approves it, and the organization records that it underwent human review. If something goes wrong, responsibility falls on the person who gave final approval. The agent making the judgment and the agent bearing responsibility have become separate. Watching this process, I sometimes think that human oversight is becoming less a mechanism for fulfilling responsibility than a procedure for preserving its outward appearance.

The same way of thinking is easy to find in corporate AI strategy today. Many organizations talk about AI transformation and set goals to convert existing work to AI. Yet on closer inspection, the fact that AI was applied often takes priority over the problem it was meant to solve. The vocabulary converges across organizations: LLM-based agents, multimodal platforms, workflow automation, and generative-AI services recur in nearly every industry. The differences between domains consequently appear to fade at the surface. Whether the industry is broadcasting, finance, manufacturing, or education, the demo can look like the same models wrapped in similar interfaces. Once a system enters real operation, however, those differences become vivid again. A translation model producing a fluent English sentence and a broadcaster reliably producing subtitles for international sale are entirely different problems. The latter requires understanding the scene, the relationships among speakers, consistent handling of proper nouns and expressions, and consideration of how mistranslation could affect both the content's meaning and the sales process.

Likewise, making a face look clean in one image is different from preserving a person's identity, skin texture, color, and frame-to-frame consistency in real broadcast footage. A result that looks plausible in a still image may reveal subtle jitter, color shifts, and artifacts in video. Nor is naturally filling a single frame after removing subtitles the same problem as consistently reconstructing moving backgrounds and objects across a long video. AI has not erased the differences among domains. It merely hides them for a moment in demos; in production, they reappear in more expensive forms. As foundation models become more alike, differentiation moves away from the model itself and toward data, evaluation, operational constraints, and the cost of failure. What ultimately matters is not which model was used, but knowing when that model fails in a particular environment and what the failure means.

This led me to think that the role of a researcher must shift from producing answers to discovering problems. If AI can generate code and documents and take over part of an experiment, the most important work left to people is deciding what to define as the problem. It is judging which assumptions distort reality, what should count as success, and which results should make us abandon the current approach.

But this raises another problem: what if raising the right problem slows down the business?

Analyzing a technical problem takes time. We have to collect real failures, separate their conditions, and verify their causes. By contrast, even a loosely assembled AI service produces something visible if it is released quickly. It generates news, material for internal reports, and reactions on LinkedIn and other social platforms. Technical maturity is hard to explain, but launches, views, articles, and partnerships are easy to express as numbers. A deep challenge produces a complicated and uncomfortable story. It reveals that current achievements may be smaller than expected, calls for further validation, and adds unresolved risks to the list. The better we become at finding problems, the more work there is to do, the later the release date becomes, and the more complicated the simple narrative management wants to present to the outside world grows. Moving closer to the truth can appear to obstruct the business.

I do not mean that a service should never launch until every technical problem has been solved. A company is not a research institute with unlimited time and resources. In many cases, releasing an imperfect product first and observing users' responses is more rational. A researcher who tries to eliminate every risk in the name of technical purity can also become someone who prevents execution. The important distinction is between releasing an imperfect product as an experiment and declaring it a success while concealing its imperfections. The former is an attempt to reduce uncertainty: operate within a limited scope to learn what works and what does not. The latter leaves uncertainty unresolved and covers it with the language of achievement.

Hype is not purely bad in itself. Attention creates budgets, gathers people, and moves organizational inertia. A good narrative gives a technology room to grow. The problem begins when hype stops being a means to build better technology and becomes a substitute for technical evidence. The existence of a user response is transformed into proof that a service operates reliably; attention on social media is presented as though it demonstrates a sustainable business.

In this environment, deep expertise is overlooked not simply because executives do not understand technology. There is a gap between what an organization rewards and what a researcher considers important. Organizations often reward visibility over technical quality. A launched service leaves behind a screen, articles, and presentations, and it is easy to explain who built what. By contrast, the achievements of someone who finds a problem in advance are hard to see. They may have prevented the adoption of a bad model, a mistranslation incident, or an operational outage by changing the architecture in time, yet all that remains is the fact that nothing happened. A launch is an event; prevention is a non-event. It is difficult to prove why a problem did not occur, and rarely is it recorded whose judgment prevented it.

Another problem is that the rewards of achievement and the costs of technical debt go to different people. The person who launches a service now receives today's evaluation and attention. The future maintenance burden, data contamination, broken promises, and loss of trust created by a rushed architecture fall to later practitioners and operations teams. A decision that is irrational from the organization's overall perspective can be perfectly rational within the tenure and evaluation criteria of a particular decision-maker. A researcher who raises a problem then appears not as someone helping the organization achieve its goals, but as someone preventing the organization from calling its preferred outcome a success. This is not because the criticism is wrong; it is because it conflicts with the objective the organization is optimizing.

Where, then, should researchers go amid all the jobs created by the AI boom? Good companies and positions that value depth, respect long-term research, and care about technical trust still exist. They are not numerous, however, and not everyone within them can be a core decision-maker. An organization needs far fewer people to make final judgments than it needs to execute them. The AI market may be growing toward a need for many people who know how to use AI, rather than a need for many AI researchers. Most new jobs will probably be less about studying frontier models than about connecting existing models to work, assembling services around them, and introducing them into organizations. Inevitably, far more people will apply, integrate, sell, and operate models than research the models themselves.

The ability to adopt new technology quickly is certainly important in that process. People with steep learning curves have an advantage in fast-changing environments. But learning quickly and understanding deeply are not substitutes for one another. If we compete only on the ability to use a new API or framework quickly, that ability resets every time the next tool arrives. Once everyone can learn tools at a similar pace, differentiation disappears just as quickly. Deep knowledge reveals itself when abstractions break: when a model behaves unexpectedly, when the data fails to reflect reality, when public benchmarks diverge from production results, or when cost, latency, and security problems arrive together. As AI takes over implementation, technical depth may become not less necessary but more important, because it is needed to judge when AI-generated results are wrong.

Believing that depth matters is not enough, however. Nor can we expect organizations and markets to recognize it automatically. Invisible expertise is often treated as if it did not exist. We must turn depth into a form that can be used in actual decisions. It is not enough to stop at saying that a model has a problem. We must show the conditions under which it occurs, the cost it creates for users and the business, what should change in the current decision, the smallest experiment that could test it, and which outcomes should tell us to continue or stop.

Researchers must move from people who state problems to people who design their costs and options. Rather than saying that something simply cannot be done, they should explain what can be done now, where the danger begins, and what must be checked before moving to the next stage. Their role is to translate technical concerns into the language of business and business demands into technical questions that can be tested.

Influence within an organization does not arise from insight alone. It grows stronger with ownership of data and systems, metrics and outcomes. Someone who builds translation evaluation data, defines error types, and owns release criteria and monitoring has more influence than someone who merely points out that translation quality is poor. Reproducing the conditions that cause video artifacts and designing quality standards and failure responses matter more than simply saying that artifacts exist. We need to own concrete systems that require judgment. That means understanding how to build the model, how data enters, how results are used, and where failures along that path become our responsibility.

If the organization still refuses to listen after all that, the story changes. If management continues to choose visibility and hype even after the conditions and costs of a problem, solutions, and alternatives have been explained clearly, the issue is no longer the researcher's communication skills. It is a conflict between different views of business and different time horizons. One side sees technical trust and accumulated capability as long-term business value; the other values current attention and evaluation more highly. Neither is a purely technical judgment. In the end, it is a choice about whose benefit, at what point in time, should count as corporate value.

It is nearly impossible for one person to change an entire organization's reward structure. Trying to correct every bad judgment quickly leads to burnout. I therefore believe we have to distinguish the organization's reward function from the reward function of our own careers. At a company, releases and schedules, user response, and business performance matter. A researcher who insists on depth while ignoring them will find it difficult to remain in the organization for long. I, too, have to produce results, collaborate with others, and explain what I know in language they can understand.

But I should not mistake only what the company rewards today for my own growth. At the end of a project, what matters alongside what I launched is which problems I understand more deeply than before. I should ask whether I have taken responsibility for a system from beginning to end, reproduced a failure and identified its cause, left behind reusable data and evaluation systems, and developed expertise that survives even when the name of the model changes.

In a good organization, company performance and personal growth move together. Building a service also builds depth, and solving a business problem generates new technical questions. In a poor environment, only the organization's ledger grows. You rapidly make demos, attach new models, and prepare presentation decks, but what remains for you is merely experience following the trend of a particular moment. That experience is necessary at first. It can teach you how to implement technology quickly, which messages move an organization, and which problems arise while deploying a real service. Yet if similar projects keep repeating without expanding the scope of your judgment and ownership, it may be repetition rather than experience.

I need to ask myself periodically: will what I am learning now remain when the next model appears? At the end of a project, are we accumulating data, evaluation criteria, and knowledge of the system? Am I being entrusted with increasingly important judgments, or merely being asked to build more convincing demos in less time? Will I retain accomplishments and capabilities that I can explain outside this organization?

Making depth visible is ultimately my responsibility as well. The market does not automatically reward truth, and accurate judgment does not always defeat a simpler, better-packaged narrative. Technical theater may, in fact, be rewarded for much longer than expected. Persuasion is part of a researcher's job. The limit is selling confidence larger than the evidence. A researcher with depth should be able to demonstrate it through reproducible experiments, evaluation data, failure cases, operational metrics, and records of technical decisions. The record should show how that understanding led to better judgments and outcomes.

Not every problem needs to be fought with equal intensity. A reversible, low-cost decision can be tried first even if it is imperfect. If an uncertain issue can be measured, we can run a limited experiment and collect data. Instead of letting a researcher's intuition and an executive's confidence collide endlessly in a meeting room, it is better to define results that will reveal whose hypothesis was wrong.

Irreversible, high-loss decisions are different. If I am asked to promise unverified performance externally as fact, to deploy without adequate review where mistranslation and distortion can affect real content and users, or to stake my name and expertise on results I have not verified, I have to raise the problem clearly. When necessary, I should document the basis of the judgment and the boundaries of responsibility. We cannot prevent every form of technical debt, but we do not have to pledge our professional credibility as collateral for someone else's bad judgment.

If these situations are recurring rather than temporary, leaving must also be considered. Some organizations may hide negative experimental results, demand that unverified performance be presented publicly as fact, and make the person who discovers a problem bear the responsibility while the person who created it takes the credit. Others may change the metric instead of improving the model when evaluation results are poor, defer technical debt to the next quarter without end, and fail to accumulate data, evaluation systems, or infrastructure despite repeating project after project.

In such an environment, it is dangerous to make changing the organization your personal mission. An individual's career may last longer than one organization's AI strategy. Conversely, a person's most important career years may be too short to wait for the technical theater to end and the market to recognize what is real. Leaving can preserve one's professional standards and long-term options. I believe a researcher's final asset is the ability to keep personal judgment independent of a broken reward structure, rather than a title received from a particular company.

This is why I have lately tried not to cling as tightly to the title of researcher. Positions devoted purely to research may become fewer. Research will remain an independent function in a small number of frontier organizations, but in most workplaces it will seep into engineering, product development, and operations. For me, preserving the way of doing research matters more than preserving the title. That means forming hypotheses, establishing falsifiable conditions, testing them against real data, refusing to hide failed results, and stating clearly what we do not know.

I will continue to use AI actively. There is no reason to reject the help it offers while I read papers, write code, design experiments, and organize documents. I do want to distinguish the delegation of execution from the delegation of judgment. The more output AI produces, the more carefully I must judge what to believe. The faster it writes code, the more I need to understand the assumptions beneath that code. The more plausible its explanations become, the more I must distinguish natural prose from factual accuracy. The more AI evaluates papers, the more I must ask whether its evaluation merely imitates past preferences or recognizes the value of new knowledge.

AI has reduced the cost of generation, but it has not reduced the cost of verifying what is correct or accepting responsibility for the result. More output means more to verify. Generation has become cheap; judgment has become expensive. A researcher's value therefore cannot be explained only by the ability to produce more. What matters is the ability to discover problems others missed, turn vague discomfort into testable questions, and explain which results can be trusted. Accurately stating what a model cannot do is as necessary as demonstrating what it can.

Discovering a problem is only the start. We have to describe its cost and available choices, produce evidence in a form an organization can use, and, if possible, own the system that solves it. If an organization has no need for judgment even when presented with precise evidence, we must accumulate enough expertise and options to leave it.

I still believe that we can compete through research by finding problems the average answer cannot see, turning uncertainty into a form we can handle, and taking responsibility for our own judgment. That is a different contest from writing more papers than AI or producing code faster. AI can take over many of the procedures of research, but it cannot automatically decide what to doubt and what to take responsibility for. Researchers may no longer be people who monopolize answers. Instead, they must become people who verify what society and organizations can safely believe.

Perhaps the thing I need to protect in the future is less the position of researcher than the habit of not accepting things too easily: embracing rapidly changing technology without surrendering judgment to its speed, understanding the narrative an organization wants without crossing the boundary of fact, and producing results today while accumulating knowledge and systems that remain tomorrow. It will not be easy. Speed will continue to be rewarded over depth, and people who discover problems will repeatedly be treated as though they created them. At times, I may wonder whether I am being overly cautious or failing to adapt to change.

Even so, I want to remember one thing: research determines which answers can be believed. Living as a researcher means being prepared to test one's own fallible judgment and accept the result. As we enter an age in which AI can make almost anything, I want to describe my work less by what I made than by what I questioned to the end and what I took responsibility for.