In my last post I mentioned the case of my having someone’s private and personal mobile phone number, and being asked for that number by someone else. There is a trust issue there, I should consider myself a steward of that information. I am not at liberty to give that information away willy-nilly to anyone who asks.
Stewardship of data that “belongs” to more than one person can be tricky. That’s a small part of the reason why “right to be forgotten” isn’t easy to implement. People had to work out what happens when someone has a photograph of multiple people, taken with their consent, and then one of those people wants to exercise the right to be forgotten. Delete the photo? It doesn’t “belong” to the deletion requester. If I remember correctly, the direction of travel then became “untag the person but leave the photo alone”. Maybe someone will come up with the idea of AI-based airbrushing of such photographs, heaven forfend. The topic of data pertaining to more than one person is not new, work has been done for quite some time to handle this.
Determination of purpose is also interesting. Network tracing during COVID became fraught with data protection danger; the situation was even more complex with passengers travelling with or sitting near an Ebola carrier. There’s a lot of law already in place that restricts data use using purpose as a filter.
Then there’s the legal aspect to it. What records have to be preserved? What can be amended? What can be deleted? These are not new questions, there’s decades of work already completed in dealing with such questions, and tools and processes exist to do this correctly, some automated, some manual. But it’s not new either.
AI agents can do magic because they are permitted to get to stuff quickly that humans can’t get to easily, and can look at a lot more stuff than humans can. incidentally, when it comes to how efficient these agents are, my starting point is that the human brain consumes 15-20 watts. I won’t even get into the quantum of other resources used by the human body. Suffice it to say that 20w is a useful benchmark.
The magic costs permissions. Permissions that come at a price. The fines for dealing with data protection failures aren’t trivial. I don’t even know what the fines are for not complying with legal disclosure and records retention.
It’s even more complicated. When looking at building data businesses within large enterprise, one of the first things I had to consider was title. What dots do we own?… not always an easy question to answer in an enterprise context, with multiple customers, institutions and partners involved.
You can’t give permission to an agent to access data you don’t fully own. Frontier models may have acquired data in all kinds of ways, seeking to justify such disregard for title using some questionable fair use argument or the other, flying the “because China” or “because cancer” pennants. The bills will arrive. Particularly when agents take off. Product liability issues will probably keep a generation of lawyers going for a while.
I’m not an AI doomer. As with my earlier post, I’m optimistic about the future. But the disregard for protections evolved in many markets over time: copyright, intellectual property, privacy and confidentiality, data protection, even listing rules … this disregard comes at a price.
The Grateful Dead fan in me has always been convinced about open source software in general. A very high proportion of the software in use in enterprise is open source, whether we want to admit it or not. When it comes to defence the ratio is not as high but it remains non-trivial. There will always be attempts to push back on open source (and in the case of frontier models open weights and open access to training data as well). Which leads me on to open data, something else we have to think about. Again this is not new, people have been working out what to do for years, and there are many good examples of open data in use.
Why do I bring open data up? Because there are many pockets of good open data, all running the risk of being contaminated, poisoned or even deleted forever by indisciplined use of agents. Firms rely on access to such data.
This isn’t about SALT talks for frontier models and the need to slow down development or issues of that ilk, whether it’s to save humankind or the universe or national security or the odd IPO or two or three. I’ll leave that to the experts.
This is about the use of agents in normal life. The role of trust and permissions in the context of agents in normal everyday life, whether enterprise or consumer. What we need to know about each agent. What datasets are touched. How permission was obtained, recorded and audited. How restrictions were tested and enforced. How title was proven. How records were appropriately dealt with, in terms of retention and deletion. How disciplines were fashioned, implemented, improved.
How trust was established. Trust.
Permissioned access to data would be a start.
For my sins I spent a few years running “death march” projects and programs across large enterprises.
Y2K was one of them. It was a problem directly caused by erstwhile disciplines, built up over decades, scattered to the four winds as democratisation of access to tools –both acquisition as well as usage — hit the enterprise world.
Firms didn’t know what software they had. Who had provided it to them. Who maintained them. What processes ran on “end user computing” and “shadow IT”. What kit they had and where. What the systems did. When, where and how. The list is endless.
It took the Euro, Y2K and the explosion of the web and mobile to catalyse people into action. Build inventories. Rebuild control processes. Build inspection and verification tools.
The best estimates I saw in those days was that around 7-8% of Y2K expense was on remediation and testing of code. The rest? Building back disciplines that were necessary and yet had been eradicated.
As agents proliferate, we have the opportunity to relive that joy. To bring back disciplines that have disappeared.
The pedigree of the agent. Datasets permitted for access, and by whom. Proof of title. Consents. “Restricted covenants”, no-do activities and no-go areas. How long the agent is licensed for use. By whom. Activity and audit logs. How long everything would be retained. Enforcement policies. This is not an exhaustive list, just a soupçon of what’s involved.
Otherwise? It’s not doomsday. Not armageddon. But the fragnant aroma of product liability lawsuits in part of the market, and legal breaches in the collection, use, preservation and deletion of data.
We can prevent this being a re-run of Y2K.
What we can’t do is to prevent a re-run of the One-Click moment. People have to trust the new world. Which will need education. Tooling. The establishment of recourse.
Clever people can work on if and whether and how to slow the frontier models down.
Me? I’m more interested in ensuring that my family and friends and colleagues and community continue to have access to things that make their life a little easier. Tools that support their health, their mental state, their education, their relaxation, perhaps even their hard-earned wealth.
We don’t need stories of how someone’s photographs and memories disappeared by agent accident. We don’t need stories of why critical operations are held up because key files are missing. We don’t need tales of bank accounts being emptied by agent misbehaviour. I’m not bothering to elaborate the cyber risk, everyone knows that. I would rather think about the unintended consequences of helter-skelter stampedes towards miasmas of personal and enterprise and global agents without sorting out permissions and recourse first.
May be a while before I continue this thread. I’ve started having some really interesting conversations, got some interesting pointers as to what to read, what to refer to, and will start planning a few face to face sessions with people who’ve got in touch with me.
Civil discourse. May it live long.