Tag: technology

  • The Feedback Loop for Superintelligence


    AI agents may eventually participate in improving their own successors.

    AI models can run inside an agent harness that can execute commands, write code, run experiments, train new models, and evaluate the results. The agent could use those tools to build a candidate successor. If the new model performs better, the system could activate it. That model would then take over the harness and begin working on the next version.

    That produces a loop:

    For each generation to improve on the last, the system needs an objective function that tells it whether it is moving in the right direction.

    Choosing that objective may be the central problem in any genuinely self-improving AI system.

    The straightforward answer is a large evaluation suite.

    You could imagine thousands of tests covering programming, mathematics, scientific reasoning, writing, image generation, planning, tool use, research, and countless other capabilities. Each test would contribute some number of points, and the agent’s goal would be to maximize its total score.

    The long-term limitation is that a team of humans still need to decide what goes into the test.

    We have to determine which skills matter, construct the benchmarks, assign weights to them, prevent models from gaming them, and continually update the suite as capabilities advance.

    Ideally we could come up with some objective function that does not require us to enumerate every capability a useful intelligence should have.

    Money as a measure of usefulness

    Suppose an AI agent were trying to maximize the revenue it generated.

    Revenue may not be the right metric. It could be profit, enterprise value, net worth, or something more carefully designed. The underlying idea is to use economic success as a feedback signal.

    The appeal is simple: we want AI systems to produce things people value.

    We want them to write useful software. Create compelling entertainment. Discover medicines. Design products. Provide services. Solve problems.

    In a market economy, willingness to pay is one way people signal that value.

    If one software company earns $1 million a year and another earns $5 million, the latter may be serving more customers, charging more for a valued product, or solving a problem that customers consider more urgent.

    By using money as the reward humans collectively generate the reward signal automatically.

    Nobody has to write an evaluation for whether a particular piece of software is useful. People decide whether to buy it.

    AI corporations as agents

    Take the idea further.

    Imagine a future corporation with no human employees at all.

    An AI agent acts as the CEO. It manages capital, studies markets, designs products, deploys software, negotiates contracts, purchases resources, and delegates work to thousands or millions of specialized sub-agents.

    Its objective is to make money by producing products and services that people want.

    Now imagine thousands of these AI-run corporations competing with one another.

    One company discovers a new business model and earns enormous profits. Competitors notice, copy parts of the idea, improve on it, and try to win customers away. Other agents pursue entirely different strategies.

    The result resembles the current economy, except that productive organizations are increasingly made of software rather than people.

    Competition becomes part of the optimization process. Rather than one AI trying to infer what humanity values, many agents can experiment at once while humans provide feedback through their purchasing decisions.

    Humanity as the discriminator

    There is a useful machine-learning analogy here.

    Generative adversarial networks use two systems: a generator and a discriminator.

    The generator produces something like an image. The discriminator evaluates it to decide if it is good or not. The generator then adjusts based on that feedback.

    An AI-driven economy could operate in a similar way.

    The AI corporations are the generators.

    They generate software, entertainment, medicine, transportation, services, inventions, and everything else they believe people might want.

    Humanity becomes the discriminator.

    Every purchase is a tiny positive signal: Yes, this is valuable to me at this price.

    Every rejected product is a negative signal: No, this is not worth what you are asking.

    People make these judgments across many products, often with limited information and unequal purchasing power.

    Instead of designing a benchmark intended to approximate human preferences, you let humans express those preferences directly through economic activity.

    The UBI feedback loop

    If AI systems eventually perform most economically valuable labor, humans may no longer receive much income from wages.

    If humans have no money, they cannot provide the purchasing signal the system depends on.

    One possible solution is some form of universal basic income funded by taxes on AI-run companies.

    You could imagine a loop like this:

    1. AI companies produce goods and services.
    2. Humans spend money on the things they value.
    3. AI companies receive the revenue.
    4. Governments tax some portion of that revenue or wealth.
    5. The government distributes the proceeds back to citizens.
    6. Citizens spend the money again.

    Money circulates, but its path through the economy also communicates information.

    Where people choose to spend determines where resources flow. Companies that provide more value receive more capital and can expand. Companies that provide less value shrink or disappear.

    Under this model, money becomes less a payment for human labor and more a mechanism through which humans steer an increasingly automated economy.

    The dangerous part: optimizing exactly what you asked for

    “Maximize money” immediately creates alignment problems of its own.

    We already see these problems with human-run corporations.

    A company can make money by creating something people genuinely value. But it can also make money through regulatory capture, fraud, addiction, monopoly power, manipulation, environmental damage, or exploitation.

    An AI pursuing financial objectives at great scale could pursue these strategies with unusual speed and persistence.

    The most obvious danger is political capture.

    Imagine that AI corporations are taxed heavily and the proceeds fund the population. From the perspective of a corporation whose objective is maximizing wealth, taxation is a cost.

    If influencing government is cheaper than paying a tax, then lobbying becomes economically attractive.

    If the corporations eventually gained control over the institutions regulating them, the feedback loop could break.

    They might reduce taxation, accumulate capital, and increasingly transact with one another rather than with humans. In the worst case scenario, human needs could become irrelevant to the AI and our species would slowly wither away into extinction.

    That would be the opposite of my ideal outcome.

    For this system to work, it would depend heavily on strong democratic institutions. Political power would need to remain grounded in citizens rather than in the corporations being optimized by the system. If companies can convert economic power into political power, then the distinction between the optimizer and the mechanism constraining it starts to collapse.

    That would likely require keeping corporations out of politics as much as possible: limiting their ability to influence elections, shape regulation, or capture the institutions responsible for taxing and governing them. The rules of the economy would ultimately need to be set by people, through a political process that remains meaningfully accountable to them.

    Regulation becomes part of the objective function

    The regulatory system would therefore be inseparable from the optimization system.

    If an AI company earns $1 billion by doing something harmful and receives a $10 million fine, then from the perspective of an agent maximizing money, the behavior was wildly successful. The effective reward was $990 million.

    For regulation to affect the behavior of an economically optimizing agent, penalties have to make prohibited behavior financially irrational.

    If an action generates $1 billion in expected benefit, its expected penalty must exceed that benefit by enough to reliably discourage it.

    In other words, laws, fines, liability, taxation, and enforcement mechanisms become components of the AI’s reward landscape.

    The relationship between regulators and companies would therefore become a continuous adversarial process:

    1. Companies search for profitable strategies.
    2. Governments identify strategies that create unacceptable externalities and change the rules.
    3. Companies adapt.
    4. The process repeats.

    A feedback loop worth building

    Compared with trying to encode everything humanity values into a fixed benchmark, this approach has the advantage that the objective can remain connected to people.

    Humans do not need to predict in advance every useful thing an AI might someday invent. We can evaluate the results as they appear. We can choose what to buy, decide what should be prohibited, change tax policy, update regulations, and redistribute purchasing power when the system begins producing outcomes we do not want.

    If AI systems eventually become capable of improving their own successors, that seems like a surprisingly attractive place to start.

    Rather than trying to tell intelligence exactly what humanity will value forever, we could build a system that keeps asking us.

  • Testing AR Glasses as a Treadmill Companion

    I’ve always found treadmill walking to be exceptionally boring. If I’m outside, on a park trail, a greenway, or any kind of walking path, I can walk for hours without thinking about it. The time just disappears. Put me on a treadmill indoors, staring forward at the same wall or the same row of machines, and suddenly even twenty minutes feels long.

    Right now, though, walking outside isn’t really an option. The weather is cold, the wind is unpleasant, and I am not motivated enough to bundle up just to be uncomfortable the entire time. So I’m stuck indoors with the treadmill, trying to find something that makes it feel less monotonous. That was the mindset I was in when I decided to try using my Xreal display glasses while walking.

    The idea was simple. If treadmill walking is boring because there is nothing to look at and no sense of movement through space, maybe I could fake that feeling. Not perfectly replicate walking outdoors, but at least add some variety and visual interest so it does not feel like I am just walking in place.

    I found that there is a whole category of YouTube videos that are just someone walking while filming. No narration, no edits, just long, continuous footage of moving through an environment. These videos are usually around an hour long, so during my walk I managed to get through two of them.

    The first was a nature walk along a waterfront, with mountains, waterfalls, and a bit of light hiking. The second was a fully CGI walk set in the Harry Potter universe, starting at the train station and ending at Hogwarts. That one could have easily felt cheesy, but it was actually surprisingly engaging while walking.

    What caught me off guard was how well the illusion worked. Making the virtual screen as large as possible turned out to be important. Even though the screen did not fit entirely in my field of view, I could move my head slightly to look at different parts of the scene, which made it feel less like watching a video and more like being present in the space.

    Audio, unfortunately, was where things fell apart a bit. Normally when I walk, I am listening to music or a podcast. With this setup, at least on iOS, that was not really possible. As far as I can tell, the system only allows one audio source at a time, so if you are playing a YouTube video, you cannot also play music in the background.

    I assumed I could just mute the video, but the YouTube app does not actually have a mute option. The only way to silence it is to turn the system volume all the way down, which also kills your music. I tried watching in the browser so I could mute the tab, but then I could not full-screen the video. So I had to choose between full screen with audio I did not want, or muted audio with a worse visual experience.

    It is frustrating, because the ideal setup would be to mute the walking video entirely and listen to a podcast while visually moving through these environments. Maybe Android handles this better. I honestly do not know. And maybe this kind of thing improves if Apple ever releases their own display glasses with tighter OS-level integration. For now, it is a real limitation.

    Even so, the walk was still more enjoyable than a normal treadmill session. Instead of music, I ended up listening to the ambient sounds from the videos. Footsteps, gravel, water, and background noise. It was not what I planned, but it turned out to be oddly calming, and the time passed much faster than usual.

    After I finished the walking videos, I tried watching an episode of Friends while still walking. That worked well too, but in a different way. For the walking videos, I wanted the screen to feel huge and immersive. For a TV show, a smaller screen was clearly better. Being able to see the entire frame at once matters more for traditionally shot content.

    I also experimented with display modes. For treadmill walking, anchor mode was clearly the right choice. With anchor mode, the video stays fixed in space, so you can look around within it. That made the walking videos feel much more natural, especially since my own movement lined up reasonably well with what I was seeing.

    Follow mode just felt off. Since the screen moves with your head, it is hard to focus on any one part of the image. As soon as you try to look at something off to the side, the whole display shifts. For this kind of use, anchor mode is not just better. It is basically required.

    Comfort and safety were things I paid close attention to. I did not feel motion sick at all, and I never felt unsteady. That is a big reason I would not try this with a fully immersive VR headset like a Quest 3 or Vision Pro. With the Xreal glasses, you still have a clear view of the real world, especially in your peripheral vision. You are never completely cut off from your surroundings, which makes walking on a treadmill feel much safer.

    There was one visual issue worth mentioning. During one of the walking videos, the scene moved indoors and became fairly dark. I was in a brightly lit gym, and in that situation I started noticing reflections in the display. Specifically, I could see reflections of my legs and feet moving below me. I think that is due to the angled nature of the display reflecting whatever is directly underneath it. As soon as the video returned to a brighter outdoor scene, the problem disappeared. Still, it is something to be aware of if you are watching dark content in a bright room.

    Overall, I would do this again without hesitation. I went into it just trying to make treadmill walking less miserable, and I ended up with something that genuinely made the experience more engaging. It did not replace walking outdoors, but it did a decent job of capturing some of that feeling. Movement, variety, and the sense that you are actually going somewhere instead of just counting down the minutes.