/ tech

Philowashing: When Philosophy Legitimizes Silicon Valley

Silicon Valley hires philosophers to rethink the values of its AI. Philowashing: a sophisticated form of corporate legitimization.

Philowashing: When Philosophy Legitimizes Silicon Valley

Those of us who studied philosophy are used to being asked what it's good for. The question can take more or less friendly forms, but it almost always presupposes a genuine difficulty: unlike medicine, engineering or law, it isn't immediately obvious what place a philosopher might occupy outside teaching and research. That's why it's tempting to read what's happening at some of the world's most important tech companies as a small bit of disciplinary payback: Silicon Valley, the very ecosystem that for decades turned the engineer and the entrepreneur into near-heroic figures of our age, seems to have discovered that it needs philosophers too.

Amanda Askell, trained in philosophy in Cambridge and New York, has become one of the most visible faces of Anthropic, the company behind Claude, where she works on the character and values that should guide the behavior of its models. Iason Gabriel, who holds a doctorate in philosophy from Oxford, has for years held various roles connected to ethics, pluralism and governance at Google DeepMind. These are not isolated cases: as artificial intelligence stopped being a relatively specialized technology and began intervening in domains ranging from education and work to health, security and personal relationships, questions that until recently might have seemed fit for a university seminar took on an unexpected practical urgency. What values should an artificial intelligence system respect? How should it act when two moral principles come into conflict? Can a model be designed that's capable of coexisting with the pluralism of societies where there's no agreement about what is good, what is just, or even what is true?

Silicon Valley, the very ecosystem that for decades turned the engineer and the entrepreneur into near-heroic figures of our age, seems to have discovered that it needs philosophers too.

That these questions have made their way into artificial intelligence labs is, in principle, good news. Nor does it seem to be merely a public relations strategy: Askell, Gabriel and other researchers (like Luciano Floridi, who worked for Google on the "right to be forgotten" after an adverse ruling by the Court of Justice of the European Union) produce philosophically sophisticated work and face genuine problems, ones for which a humanities background can contribute something you don't get by simply adding more and more engineers to a team. Still, it does seem fair to us to look not so much at the quality of the answers these philosophers produce, but at the institutional conditions within which they get to formulate their questions. If a tech company hires philosophers to think about what values a system it has already decided to build should incorporate, what happens to the prior — and potentially more uncomfortable — question of whether that system should occupy the place the company intends to assign it?

Our suspicion is that we're dealing with something deeper than just another case of ethics washing (that is, cosmetic changes meant to give an ethical gloss to an action); rather, we're witnessing a kind of appropriation of our discipline in order to delimit in advance the range of problems it's allowed to weigh in on. We propose calling that mechanism philowashing ("filolavado" sounds weird!) and we think it has a precedent in the contemporary appropriation of Stoicism by the "tech bros."

From Stoicism to broicism

Over the past two decades, Marcus Aurelius, Epictetus and Seneca have gained unexpected popularity among entrepreneurs, investors and figures associated with tech culture. Tim Ferriss helped turn Stoicism into a standard reference within the world of productivity and personal development; Ryan Holiday built a wildly successful ecosystem of books, courses and conferences around that tradition; and Marcus Aurelius's Meditations ended up sharing shelf space in corporate libraries with manuals on leadership, growth hacking and performance optimization. The phenomenon is curious, though not necessarily objectionable: someone arriving at Epictetus in search of tools to get through a professional crisis doesn't invalidate either the reading or the philosophy they discover along the way.

What's significant is the selection that appropriation performs. Ancient Stoicism wasn't a collection of techniques for boosting productivity but a considerably more ambitious conception of virtue, duty, life in common and the relationship between what depends on us and what we can't control. In its translation into contemporary corporate culture, however, much of that philosophical architecture gets set aside, and what survives above all are the ideas that can be turned into techniques for managing the self: tolerating adversity, controlling your own reactions, focusing on what you can actually affect and building resilience in the face of external circumstances you can't change.

If a tech company hires philosophers to think about what values a system it has already decided to build should incorporate, what happens to the prior, and potentially more uncomfortable, question?

The result has sometimes been called broicism, or broicism, and the term helps identify a relevant shift. If a workplace demands constant availability, produces anxiety or makes the boundary between work and rest increasingly porous, the reformulated Stoic question is no longer directed at how that environment is organized but at the individual's capacity to inhabit it successfully. Philosophy then supplies tools for intervening on the subject while the conditions producing their distress remain off the table.

Philosophers buy in at Silicon Valley

Amanda Askell's case is especially useful for understanding this shift. Part of her work at Anthropic consists in thinking about how Claude should behave: what traits its "personality" should have, what values it should express, how it should react to morally complex situations and to what extent it's appropriate for an AI assistant to try to influence its users' beliefs or decisions. Anthropic's Constitutional AI project carries that concern into the training architecture itself: instead of merely correcting undesirable responses one by one, the aim is to guide the model through an explicit set of principles.

Philosophy can play a meaningful role in those discussions. Asking what it means to respect a user's autonomy, how to avoid paternalism, or how an artificial intelligence can act in contexts where incompatible value systems coexist isn't a coat of humanistic varnish added at the end of a technical process; these are genuine normative problems. Something similar goes for Iason Gabriel's work on pluralism and alignment at DeepMind. But both the question of Claude's personality and the discussion of what values a large language model should incorporate start from a prior decision there seems to be no way to intervene in, one tied to the very project of that model, its scale, its business model and how it will come to occupy certain positions in social life. Philosophical reflection begins once much of the problem's institutional architecture has already been settled. What gets discussed is how the system should behave, but not necessarily what kinds of decisions should be handed over to it, to whom it should answer, or whether there are domains where its use should be ruled out regardless of how well it performs.

The algorithm that wasn't wrong

In late July 2019, Guillermo Federico Ibarrola was arrested at Retiro station after the fugitive facial recognition system used by the Buenos Aires City government identified him as a man wanted for an aggravated robbery committed in Bahía Blanca. He spent six days in custody until it was established that he had no connection to the crime and had never even been to that city.

The case could be filed without much trouble into the bulging archive of algorithmic errors we've piled up in recent years, except for one detail that makes it philosophically more interesting: the algorithm hadn't made a mistake. It didn't confuse two similar faces or produce a false positive due to a statistical glitch. The arrest warrant was for Guillermo Walter Ibarrola, but someone had linked Guillermo Federico's ID number to that warrant. The system correctly processed the information it received, compared the right face against the wrong database and made it possible to locate the wrong person with admirable efficiency. It was, if you like, a case of surgical precision applied to faulty data.

The key point is that no improvement to the facial recognition algorithm would have solved the problem. We could imagine a system ten times more accurate, free of demographic bias and subjected to the best available audits, and Guillermo Federico Ibarrola would still have ended up in custody. The error lay elsewhere: in the data feeding the system, in the institutional chain that had produced it and, above all, in the place the technological recommendation occupied within a process involving police officers, public officials and court personnel.

What's more, during those six days there were human beings taking part in the decision. That matters because one of the most common responses to the risks of automation consists precisely in guaranteeing that there's always a human in the loop, a person who keeps the final say. The Ibarrola case shows why that formula, while reasonable, says considerably less than it seems to. The mere presence of a human being at some point in the circuit doesn't guarantee effective oversight if that person doesn't receive the relevant information, lacks the time needed to review it, has no incentive to contradict the result or, quite simply, has no authority to change it.

The mere presence of a human being at some point in the circuit doesn't guarantee effective oversight if that person doesn't receive the relevant information or, quite simply, has no authority to change it.

In other words, "keeping a human in the loop" can describe completely different institutional setups. A doctor who receives a suggested diagnosis and can ignore it after reviewing the tests is not in the same position as an employee who has to process hundreds of automatic recommendations a day and justify in writing every time they depart from one. In both cases there is formally human oversight; in only one of them is it plausible to claim that the person's judgment retains a substantive function. Oversight, then, isn't a remedy tacked on at the end to make up for what the machine still can't do. It's a parameter someone configures in advance: how much information the reviewer receives, how much time they have, in which direction they're allowed to deviate and what consequences they face when they do.

Keeping the form, losing the function

The Spanish VioGén system lets us look at the problem from another angle. Used to assess risk in cases of gender violence, it assigns levels that can influence the protective measures a woman receives after filing a complaint. The automated assessment doesn't eliminate human intervention: officers can modify the risk level suggested by the system. That capacity, however, runs in one particular direction. They can raise it, but not lower it.

Here again it's important to resist the temptation to present this design as evidence of technological perversity. There are good reasons to adopt a precautionary standard in situations of gender violence, where underestimating risk can have irreversible consequences. Maybe that asymmetry is exactly what we want. But if so, what matters is noticing that the distribution of authority between person and system wasn't discovered by the algorithm, nor does it follow from its statistical accuracy: someone decided what the human officer could do, what the tool could do and in which direction a disagreement between the two should be resolved.

This point lets us state more precisely a problem that tends to get obscured in the contemporary debate about artificial intelligence. The public conversation is dominated by a question about capabilities: we want to know what AI can do, whether it diagnoses better than a doctor, codes better than a developer, grades better than a teacher, or detects patterns a police officer could never spot. It's a legitimate question, but the answer doesn't tell us what we should allow it to do. Knowing that a system predicts criminal recidivism with a certain degree of accuracy doesn't settle whether we want that prediction to play a part in a sentence; showing that a model identifies a disease better than a doctor doesn't determine what authority it should have over treatment; confirming that an artificial intelligence can grade exams with fewer errors than a teacher doesn't resolve whether we want to organize education around that assessment.

Someone decided what the human agent could do, what the tool could do, and which way a disagreement between the two should be resolved.

Ultimately, it's about distinguishing between two questions that often get mixed together. One concerns the system's behavior: what it should answer, what it should refuse, what values it should respect, how accurately it performs a task. The other concerns its institutional place: which decisions we hand over to it, which we keep for ourselves, who can depart from its outputs and under what conditions. The first can be answered by modifying the product. The second requires deciding the limits within which that product will be allowed to operate and, in some cases, concluding that a given application shouldn't exist even if it were technically possible to build it.

A perfectly sound philosophy

Now we can return to Silicon Valley's philosophers and get a better sense of where the philowashing might lie. The problem isn't that Askell, Gabriel or other researchers necessarily produce bad philosophy, or that they're obliged to reach conclusions favorable to their employers. In fact, the mechanism would be far less interesting if it depended on philosophers willing to say things they believe to be false. What's peculiar about this form of legitimation is that it can work even when everyone involved is working with rigor and intellectual honesty.

The reason is institutional. The question of how a model should behave is compatible with almost any decision the company has already made about where to deploy it, which clients to sell it to and what human tasks it will replace. You can pour enormous effort into making a chatbot respectful of human autonomy, pluralistic in the face of differing moral conceptions and careful with vulnerable users without thereby revisiting the working conditions of the people who produce its data, the material impact of the infrastructure needed to train it, the concentration of power its adoption generates, or the uses certain clients might make of the technology.

Philosophical work on the system's behavior also has a particularly convenient property: it can improve the product. A model that is less racist, less manipulative, more prudent and better at offering explanations is also a product that's easier to deploy and sell. There's nothing objectionable about that, just as there's nothing objectionable about building a safer car. The difficulty arises when that set of improvements is presented as evidence that the relevant normative questions have already been raised, when in fact only one particular class of them has been answered.

That's where philowashing comes in: it doesn't operate mainly by censoring uncomfortable conclusions, a strategy that's too visible and clumsy, but through an institutional division of questions. The question of internal design has budgets, teams and career paths. The question of institutional limits is in a considerably more precarious position: companies tend not to raise it because it can shrink the market for their products; regulators tend to arrive once the technologies have already been deployed; and academia, which still has spaces where it can be raised regardless of profitability, is paradoxically presented as the place philosophers have finally managed to escape from.

This is an important difference from other forms of reputational laundering. Greenwashing works when a company constructs an appearance of environmental commitment that doesn't match its practices; ethics washing, when the vocabulary of ethics provides legitimacy without substantially altering the distribution of power. Philowashing, as we understand it here, can be harder to recognize because it doesn't need to produce a false appearance. It can fund genuine philosophy, yield intellectually valuable results and even improve some products. Its legitimating effect comes from something more elementary: turning one part of the philosophical space into a stand-in for the whole and making certain questions look like <em>the</em> relevant questions about technology.

The shift happens less through falsification than through framing: a recognizable part of philosophy is preserved while the context in which some of its questions might prove politically uncomfortable disappears.

In that sense, the contemporary appropriation of philosophy shares a structure with the broicism mentioned earlier. The problem wasn't that executives misread every line of Marcus Aurelius, but that a particular selection turned a philosophy about the good life and our relationship with others into an individual technology for better tolerating an environment that remained off the table. In both cases, the shift happens less through falsification than through framing: a recognizable part of philosophy is preserved while the context in which some of its questions might prove politically uncomfortable disappears.

The question nobody owns

This matters because decisions about artificial intelligence's institutional place don't happen in the hypothetical future of an artificial general intelligence that wakes up and decides what to do with us. They're happening now, in far less spectacular ways, every time a public agency buys a facial recognition system, a company brings in a tool to evaluate workers, a hospital adopts a predictive model or a school uses an algorithm to flag at-risk students. As Javier Pallero showed in 421 when analyzing the spread of surveillance technologies, important transformations often advance tender by tender, decree by decree, through local decisions that seem too small to warrant a discussion about the conception of power they carry built into them.

The debate about artificial intelligence risks producing a similar effect. While we argue over whether large models "understand," when artificial general intelligence will arrive or whether a machine could become conscious, a multitude of less cinematic decisions quietly redistributes authority. A recommendation, a risk score, an alert, a ranked list or an automatic classification may look like modest instruments, but each one establishes who has to justify what, which outcome functions as the default and who bears the cost of departing from it.

There's also something lost in delegation that can't be expressed solely in terms of efficiency or accuracy. When a decision affecting a person is made by another person, there is at least ideally a practice of giving and asking for reasons. Those reasons may be bad, arbitrary or even false, but they can be discussed, challenged and submitted to review procedures. An algorithmic system can preserve the appearance of that practice by producing convincing explanations without those explanations actually describing the process that led to the outcome. Cosmetic explanation thus retains the form of justification while weakening its function, just as a merely nominal human oversight retains the form of control even though it has stripped the supposed supervisor of any authority.

That's why it isn't enough to demand that there always be a human in the loop. We have to ask which human, in which loop, and doing exactly what. Nor is it enough to demand that systems be explainable, because an explanation is only politically relevant if it allows us to understand, discuss and eventually modify whatever produced the decision. Preserving the institutional forms we associate with autonomy, responsibility and control is relatively easy; preserving the functions those forms are supposed to fulfill requires looking at the concrete distribution of power these technologies produce.

So what is a philosopher good for?

Maybe we should return, in the end, to the question we started with. Artificial intelligence does need philosophy, but that doesn't mean every way of bringing philosophers in is equivalent, or that the presence of an ethics department proves an organization is willing to submit its decisions to philosophical scrutiny. An institution can have excellent philosophers and, at the same time, keep some of the most important premises of its activity off the table.

The philosophical task that strikes us as most urgent, then, isn't to abandon tech companies or to automatically distrust the people working inside them. We need people capable of arguing from within about how these systems should behave, just as we need independent researchers, regulators and civil society organizations capable of raising questions no company has sufficient incentive to formulate on its own. The problem arises when we mistake the existence of the first kind of reflection for the presence of the second.

Artificial intelligence does need philosophy, but that doesn't mean every way of bringing philosophers in is equivalent.

A philosophy that thinks about technological transformation should also recover a certain calling that fascination with AI's grand dilemmas has displaced. Questions about whether a machine can think, understand or become conscious are philosophically appealing, but they shouldn't stop us from asking who produces the data it runs on, what human labor stays hidden behind automation, what resources it consumes, which agencies buy the systems, which vendors end up in positions of dependence and what real possibilities the people affected have to challenge its outputs. It's about remembering that a technology can pose ontological, epistemological, ethical and political problems all at once, and that companies will have very different reasons for funding the analysis of each one. That need to recover a material critique and to repoliticize philosophical reflection was also part of the original proposal.

Silicon Valley may have finally discovered what philosophers are good for. The question we have to ask is a slightly different one: what philosophy would want to work there for. Because a discipline whose history is shot through with the rather uncomfortable habit of examining a question's premises shouldn't settle too quickly for having secured a seat at the table. Sometimes the most important philosophical task consists precisely in asking who decided what could be discussed there.

Sumate a 421 →