About testing AI in cyber-physical settings

IEEE has an article on real-world testing of AI systems ‘in the wild’. An interesting new development in cyber-physical systems.

The article describes some of the experiments run “by Andon Labs, an AI safety company based in San Francisco that puts AI agents in charge of real-world operations and watches what happens.” These experiments take place in physical urban spaces like consumer retail shops, cafés, radio studios, etc. Here, AI agents created by e.g. Anthropic, Google, and OpenAI have to manage that specific situation. They often fail spectacularly: “Their spectacular and absurd failures have won the company plenty of attention.”

Why interesting? First, AI agents are is an interesting new direction for ‘cyber-physical systems’ to be deployed in actual urban settings.

Second, real-world testing can be seen as a kind of (corporate) storytelling about emergent proof-of-concept technologies. Such real-world tests beyond the lab allow people to start imagining these technologies in their everyday lives. I’m thinking of the work by Noortje Marres on self-driving cars.

Third, these kind of quotes bug me a lot, since it talks about human apprehensions and dislikes as some kind of barrier to be overcome, instead of speaking to inherent fallacies built into the so-called ‘impressive’ AIs: “A real store can reveal […] whether customers want to shop at an AI-run business (the early results on that last point are decidedly negative). Such social and organizational barriers may help explain why impressive AI capabilities haven’t yet translated into widespread adoption across the economy.” Come on, the problem isn’t people, the problem is tech and its makers!

Link to original article >>