Tag: superintelligence

  • The Feedback Loop for Superintelligence


    AI agents may eventually participate in improving their own successors.

    AI models can run inside an agent harness that can execute commands, write code, run experiments, train new models, and evaluate the results. The agent could use those tools to build a candidate successor. If the new model performs better, the system could activate it. That model would then take over the harness and begin working on the next version.

    That produces a loop:

    For each generation to improve on the last, the system needs an objective function that tells it whether it is moving in the right direction.

    Choosing that objective may be the central problem in any genuinely self-improving AI system.

    The straightforward answer is a large evaluation suite.

    You could imagine thousands of tests covering programming, mathematics, scientific reasoning, writing, image generation, planning, tool use, research, and countless other capabilities. Each test would contribute some number of points, and the agent’s goal would be to maximize its total score.

    The long-term limitation is that a team of humans still need to decide what goes into the test.

    We have to determine which skills matter, construct the benchmarks, assign weights to them, prevent models from gaming them, and continually update the suite as capabilities advance.

    Ideally we could come up with some objective function that does not require us to enumerate every capability a useful intelligence should have.

    Money as a measure of usefulness

    Suppose an AI agent were trying to maximize the revenue it generated.

    Revenue may not be the right metric. It could be profit, enterprise value, net worth, or something more carefully designed. The underlying idea is to use economic success as a feedback signal.

    The appeal is simple: we want AI systems to produce things people value.

    We want them to write useful software. Create compelling entertainment. Discover medicines. Design products. Provide services. Solve problems.

    In a market economy, willingness to pay is one way people signal that value.

    If one software company earns $1 million a year and another earns $5 million, the latter may be serving more customers, charging more for a valued product, or solving a problem that customers consider more urgent.

    By using money as the reward humans collectively generate the reward signal automatically.

    Nobody has to write an evaluation for whether a particular piece of software is useful. People decide whether to buy it.

    AI corporations as agents

    Take the idea further.

    Imagine a future corporation with no human employees at all.

    An AI agent acts as the CEO. It manages capital, studies markets, designs products, deploys software, negotiates contracts, purchases resources, and delegates work to thousands or millions of specialized sub-agents.

    Its objective is to make money by producing products and services that people want.

    Now imagine thousands of these AI-run corporations competing with one another.

    One company discovers a new business model and earns enormous profits. Competitors notice, copy parts of the idea, improve on it, and try to win customers away. Other agents pursue entirely different strategies.

    The result resembles the current economy, except that productive organizations are increasingly made of software rather than people.

    Competition becomes part of the optimization process. Rather than one AI trying to infer what humanity values, many agents can experiment at once while humans provide feedback through their purchasing decisions.

    Humanity as the discriminator

    There is a useful machine-learning analogy here.

    Generative adversarial networks use two systems: a generator and a discriminator.

    The generator produces something like an image. The discriminator evaluates it to decide if it is good or not. The generator then adjusts based on that feedback.

    An AI-driven economy could operate in a similar way.

    The AI corporations are the generators.

    They generate software, entertainment, medicine, transportation, services, inventions, and everything else they believe people might want.

    Humanity becomes the discriminator.

    Every purchase is a tiny positive signal: Yes, this is valuable to me at this price.

    Every rejected product is a negative signal: No, this is not worth what you are asking.

    People make these judgments across many products, often with limited information and unequal purchasing power.

    Instead of designing a benchmark intended to approximate human preferences, you let humans express those preferences directly through economic activity.

    The UBI feedback loop

    If AI systems eventually perform most economically valuable labor, humans may no longer receive much income from wages.

    If humans have no money, they cannot provide the purchasing signal the system depends on.

    One possible solution is some form of universal basic income funded by taxes on AI-run companies.

    You could imagine a loop like this:

    1. AI companies produce goods and services.
    2. Humans spend money on the things they value.
    3. AI companies receive the revenue.
    4. Governments tax some portion of that revenue or wealth.
    5. The government distributes the proceeds back to citizens.
    6. Citizens spend the money again.

    Money circulates, but its path through the economy also communicates information.

    Where people choose to spend determines where resources flow. Companies that provide more value receive more capital and can expand. Companies that provide less value shrink or disappear.

    Under this model, money becomes less a payment for human labor and more a mechanism through which humans steer an increasingly automated economy.

    The dangerous part: optimizing exactly what you asked for

    “Maximize money” immediately creates alignment problems of its own.

    We already see these problems with human-run corporations.

    A company can make money by creating something people genuinely value. But it can also make money through regulatory capture, fraud, addiction, monopoly power, manipulation, environmental damage, or exploitation.

    An AI pursuing financial objectives at great scale could pursue these strategies with unusual speed and persistence.

    The most obvious danger is political capture.

    Imagine that AI corporations are taxed heavily and the proceeds fund the population. From the perspective of a corporation whose objective is maximizing wealth, taxation is a cost.

    If influencing government is cheaper than paying a tax, then lobbying becomes economically attractive.

    If the corporations eventually gained control over the institutions regulating them, the feedback loop could break.

    They might reduce taxation, accumulate capital, and increasingly transact with one another rather than with humans. In the worst case scenario, human needs could become irrelevant to the AI and our species would slowly wither away into extinction.

    That would be the opposite of my ideal outcome.

    For this system to work, it would depend heavily on strong democratic institutions. Political power would need to remain grounded in citizens rather than in the corporations being optimized by the system. If companies can convert economic power into political power, then the distinction between the optimizer and the mechanism constraining it starts to collapse.

    That would likely require keeping corporations out of politics as much as possible: limiting their ability to influence elections, shape regulation, or capture the institutions responsible for taxing and governing them. The rules of the economy would ultimately need to be set by people, through a political process that remains meaningfully accountable to them.

    Regulation becomes part of the objective function

    The regulatory system would therefore be inseparable from the optimization system.

    If an AI company earns $1 billion by doing something harmful and receives a $10 million fine, then from the perspective of an agent maximizing money, the behavior was wildly successful. The effective reward was $990 million.

    For regulation to affect the behavior of an economically optimizing agent, penalties have to make prohibited behavior financially irrational.

    If an action generates $1 billion in expected benefit, its expected penalty must exceed that benefit by enough to reliably discourage it.

    In other words, laws, fines, liability, taxation, and enforcement mechanisms become components of the AI’s reward landscape.

    The relationship between regulators and companies would therefore become a continuous adversarial process:

    1. Companies search for profitable strategies.
    2. Governments identify strategies that create unacceptable externalities and change the rules.
    3. Companies adapt.
    4. The process repeats.

    A feedback loop worth building

    Compared with trying to encode everything humanity values into a fixed benchmark, this approach has the advantage that the objective can remain connected to people.

    Humans do not need to predict in advance every useful thing an AI might someday invent. We can evaluate the results as they appear. We can choose what to buy, decide what should be prohibited, change tax policy, update regulations, and redistribute purchasing power when the system begins producing outcomes we do not want.

    If AI systems eventually become capable of improving their own successors, that seems like a surprisingly attractive place to start.

    Rather than trying to tell intelligence exactly what humanity will value forever, we could build a system that keeps asking us.