Monday, October 09, 2023

Book excerpt: The problem is not the algorithm

(This is an excerpt from the draft of my book. Please let me know if you like it and want more of these.)

“The Algorithm,” in scare quotes, is an oft-attacked target. But this obscures more than it informs.

It creates the image of some mysterious algorithm, intelligent computers controlling our lives. It makes us feel out of control. After all, if the problem is “the algorithm”, who is to blame?

When news articles talk about lies, scams, and disinformation, they often blame some all-powerful, mysterious algorithm as the source of the troubles. Scary artificial intelligence controls what we see, they say. That grants independence, agency, and power where none exists. It shifts responsibility away from the companies and teams that create and tune these algorithms and feed them the data that causes them to do what they do.

It's wrong to blame algorithms. People are responsible for the algorithms. Teams working on these algorithms and the companies that use them in their products have complete control over the algorithms. Every day, teams make choices on tuning the algorithms and what data goes into the algorithms that change what is emphasized and what is amplified. The algorithm is nothing but a tool, a tool people can control and use any way they like.

It is important to demystify algorithms. Anyone can understand what these algorithms do and why they do it. “While the phrase ‘the algorithm’ has taken on sinister, even mythical overtones, it is, at its most basic level, a system that decides a post’s position on the news feed based on predictions about each user’s preferences and tendencies,” wrote the Washington Post, in an article “How Facebook Shapes Your Feed.” How people tune and optimize the algorithms determines “what sorts of content thrive on the world’s largest social network and what types languish.”

We are in control. We are in control because “different approaches to the algorithm can dramatically alter the categories of content that tend to flourish.” Choices that teams and companies make about how to tune wisdom of the crowd algorithms make an enormous difference for what billions of people see every day.

You can think of all the choices for tuning the algorithms as a bunch of knobs you can turn. Turn that knob to make the algorithm show some stuff more and other stuff less.

When I was working at Amazon many years ago, an important knob we thought hard about turning was how much new items were recommended. When recommending books, one choice would tend to show people more older books that they might like. Another choice we could make would show people more new releases, such as new books that came out in the last year or two. On the one hand, people are particularly unlikely to know about a new release, and new books, especially by an author or in a genre you tend to read, can be particularly interesting to hear about. On the other hand, if you go by how likely you are to buy a book, maybe the algorithm should recommend older books. Help people discover something new or maximize sales today, our team had a choice in how to tune the algorithm.

Wisdom of the crowds works by summarizing people’s opinions. Another way that people control the algorithms is through the information about what people like, buy, and find interesting and useful. For example, if many people post positive reviews of a new movie, the average review of that movie might be very high. Algorithms use those positive reviews. This movie looks popular! People who haven’t seen it yet might want to hear about it. And people who watched similar other movies, such as movies in the same genre or with the same actors, might be particularly interested in hearing about this new movie.

The algorithms summarize what people are doing. They calculate and collate what people like and don’t like. What people like determines what the algorithms recommend. The data about what people like controls the algorithms.

But that means that people can change what the algorithms do through changing the data about what it seems like people like. For example, let’s say someone wants to sell more of their cheap flashlights, and they don’t really care about the ethics of how they get more sales. So they pay for hundreds of people to rate their flashlight with a 5-star review on Amazon.

If Amazon uses those shilled 5-star reviews in their recommendation engines and search rankers, those algorithms will mistakenly believe that hundreds of people think the flashlights are great. Everyone will see and buy the terrible flashlights. The bad guys win.

If Amazon chooses to treat that data as inauthentic, faked, bought-and-paid-for, and then ignores those hundreds of paid reviews, that poorly-made flashlight is far less likely to be shown to and bought by Amazon customers. After all, most real people don’t like that cheap flashlight. The bad guys tried hard to fake being popular, but they lost in the end.

The choice of what data is used and what is discarded makes an enormous difference in what is amplified and what people are likely to see. And since wisdom of the crowd algorithms assume that each vote for what is interesting and popular is independent, the choice of what votes are considered, and whether ballot-box stuffing is allowed, makes a huge difference in what people see. Humans make these choices. The algorithms have no agency. It is people, working in teams inside companies, that make choices on how to tune algorithms and what data is used by wisdom of the crowd algorithms. Those people can choose to do things one way, or they can choose to do them another way.

“Facebook employees decide what data sources the software can draw on in making its predictions,” reported the Washington Post. “And they decide what its goals should be — that is, what measurable outcomes to maximize for, and the relative importance of each.”

Small choices by teams inside of these companies can make a big difference for what the algorithms do. “Depending on the lever, the effects of even a tiny tweak can ripple across the network,” wrote the Washington Post in another article titled “Five Points for Anger, One Point for Like”. People control the algorithms. By tuning the algorithms, teams inside Facebook are “shaping whether the news sources in your feed are reputable or sketchy, political or not, whether you saw more of your real friends or more posts from groups Facebook wanted you to join, or if what you saw would be likely to anger, bore or inspire you.”

It's hard to find the right solutions if you don't first correctly identify the problem. The problem is not the algorithm. The problem is how people optimize the algorithm. People control what the algorithms do. What wisdom of the crowd algorithms choose to show depends on the incentives people have.

(This was an excerpt from the draft of my book. Please let me know if you like it and want more.)

Saturday, October 07, 2023

Book excerpt: Metrics chasing engagement

(This is an excerpt from the draft of my book. Please let me know if you like it and want more.)

Let’s say you are in charge of building a social media website like Facebook. And you want to give your teams a goal, a target, some way to measure that what they are about to launch on the website is better than what came before.

One metric you might think of might be how much people engage with the website. You might think, every click, every like, every share, you can measure those. The more the better! We want people clicking, liking, and sharing as much as possible. Right?

So you tell all your teams, get people clicking! The more likes the better! Let’s go!

Teams are always looking for ways to optimize the metrics. Teams are constantly changing algorithms. If you tell your teams to optimize for clicks, what you will see is that soon recommender and ranker algorithms will change what they show. Up at the top of any recommendations and search results will be the posts and news predicted to get the most clicks.

Outside of the company, people will also notice and change what they do. They will say, this article I posted didn’t get much attention. But this one, wow, everyone clicked on it and reshared it. And people will create more of whatever does well on your site with the changes your team made to your algorithms.

All sounds great, right? What could go wrong?

The problem is what attracts the most clicks. What you are likely to click on are things that provoke strong emotions, such as hatred, disbelief, anger, or lust. This means what gets the most clicks are things that are lies, sensationalistic, provoking, or pornographic. The truth is boring. Posts of your Aunt Mildred’s flowers might make you happy. But they won’t get a click. But, oh yeah, that post with scurrilous lies about some dastardly other, that likely will get engagement.

Cecilia Kang and Sheera Frenkel wrote a book about Facebook, An Ugly Truth. In it, they describe the problem with how Facebook optimized its algorithms: “Over the years, the platform’s algorithms had gotten more sophisticated at identifying the material that appealed most to individual users and were prioritizing it at the top of their feeds. The News Feed operated like a finely tuned dial, sensitive to that photograph a user lingered on longest, or the article they spent the most time reading. Once it had established that the user was more likely to view a certain type of content, it fed them as much of it as possible.”

The content the algorithms fed to people, the content the algorithms chose to put on top and amplify, was not what made people content and satisfied. It was whatever would provide a click right now. And what would provide a click right now was often enraging lies.

“Engagement was 50 percent higher than in 2018 and 10 percent higher than in 2017,” wrote the author of the book The Hype Machine. “Each piece of content is scored according to our probabilities of engaging with it, across the several dozen engagement measures. Those engagement probabilities are aggregated into a single relevance score. Once the content is individually scored (Facebook’s algorithm considers about two thousand pieces of content for you every time you open your newsfeed), it’s ranked and shown in your feed in order of decreasing relevance.”

Most people will not read past the top few items in search results or on recommendations. So what is at the top is what matters most. In this case, by scoring and ordering content by likelihood of engagement, the content being amplified was the most sensationalistic content.

Once bad actors outside of Facebook discovered the weaknesses of the metrics behind the algorithms, they exploited it. From an article titled “Troll Farms Reached 140M Americans,” these are “easily exploited engagement based ranking systems … At the heart of Feed ranking, there are models that predict the probability a user will take an engagement action. These are colloquially known as P(like), P(comment), and P(share).” That is, the models use a prediction of the probability that people will like the content, the probability that they will share it, and so forth. Hao cited an internal report from Facebook that said that these “models heavily skew toward content we know to be bad.” Bad content includes hate speech, lies, and plagiarized content.

“Bad actors have learned how to easily exploit the systems,” said former Facebook data scientist Jeff Allen. “Basically, whatever score a piece of content got in the models when it was originally posted, it will likely get a similar score the second time it is posted … Bad actors can scrape … and repost … to watch it go viral all over again.” Anger, outrage, lies, and hate, all of those performed better on engagement metrics. They don’t make people satisfied. They make people more likely to leave in disgust than keep coming back. But they do make people likely to click right now.

It is by no means necessary to optimize for short-term engagement. Sarah Frier in the book No Filter describes how Instagram, in its early years, looked at what was happening at Facebook and made a different choice: “They decided the algorithm wouldn’t be formulated like the Facebook news feed, which had a goal of getting people to spend more time on Facebook … They knew where that road had led Facebook. Facebook had evolved into a mire of clickbait … whose presence exacerbated the problem of making regular people feel like they didn’t need to post. Instead Instagram trained the program to optimize for ‘number of posts made.’ The new Instagram algorithm would show people whatever posts would inspire them to create more posts.” While optimizing for the number of posts made also could have bad incentives, such as encouraging spamming, most important is considering the incentives created by the metrics you pick and questioning whether your current metrics are the best thing for the long-term of your business.

YouTube is an example of a company that picked problematic metrics years ago, but then questioned what was happening, noticed the problem, and then fixed their metrics in recent years. While researchers noted problems with YouTube’s recommender system amplifying terrible content many years ago, in recent years they have mostly concluded that YouTube no longer algorithmically amplifies — though they do still host — hate speech and other harmful content.

The problem started, as described by the authors of the book System Error, when a Vice President at YouTube “wrote an email to the YouTube executive team arguing that ‘watch time, and only watch time’ should be the objective to improve at YouTube … He equated watch time with user happiness: if a person spends hours a day watching videos on YouTube, it must reveal a preference for engaging in that activity.” The executive went on to claim, “When users spend more of their valuable time watching YouTube videos, they must perforce be happier with those videos.”

It is important to realize that YouTube is a giant optimization machine, with teams and systems targeting whatever metric it is given to maximize that metric. In the paper “Deep Neural Networks for YouTube Recommendations,” YouTube researchers describe it: “YouTube represents one of the largest scale and most sophisticated industrial recommendation systems in existence … In a live experiment, we can measure subtle changes in click-through rate, watch time, and many other metrics that measure user engagement … Our goal is to predict expected watch time given training examples that are either positive (the video impression was clicked) or negative (the impression was not clicked).”

The problem is that optimizing your recommendation algorithm for immediate watch time, which is an engagement metric, tends to show sensationalistic, scammy, and extreme content, including hate speech. As BuzzFeed reporters wrote in an article titled “We Followed YouTube’s Recommendation Algorithm Down the Rabbit Hole”: “YouTube users who turn to the platform for news and information — more than half of all users, according to the Pew Research Center — aren’t well served by its haphazard recommendation algorithm, which seems to be driven by an id that demands engagement above all else.”

The reporters described a particularly egregious case: “How many clicks through YouTube’s Up Next recommendations does it take to go from an anodyne PBS clip about the 116th United States Congress to an anti-immigrant video from a designated hate organization? Thanks to the site’s recommendation algorithm, just nine.” But the problem was not isolated to just a small number of examples. At the time, there were a “high percentage of users who say they’ve accepted suggestions from the Up Next algorithm — 81%.” The problem is that the optimization engines for their recommender algorithms. “It’s an engagement monster.”

The “algorithm decided which videos YouTube recommended that users watch next; the company said it was responsible for 70 percent of the one billion hours a day people spent on YouTube. But it had become clear that those recommendations tended to steer viewers toward videos that were hyperpartisan, divisive, misleading or downright false.” The problem was optimizing for an engagement metric like watch time.

Why does this happen? In any company, in any organization, you get what you measure. When you tell your teams to optimize for a certain metric, that they will get bonuses and be promoted if they optimize for that metric, they will optimize the hell out of that metric. As Bloomberg reporters wrote in an article titled “YouTube Executives Ignored Warnings,” “Product tells us that we want to increase this metric, then we go and increase it … Company managers failed to appreciate how [it] could backfire … The more outrageous the content, the more views.”

This problem was made substantially worse at YouTube by outright manipulation of YouTube’s wisdom of the crowd algorithms by adversaries, who effectively stuffed the ballot box for what is popular and good with votes from fake or controlled accounts. As Guardian reporters wrote, “Videos were clearly boosted by a vigorous, sustained social media campaign involving thousands of accounts controlled by political operatives, including a large number of bots … clear evidence of coordinated manipulation.”

The algorithms optimized for engagement, but they were perfectly happy to optimize for fake engagement, clicks and views from accounts that were all controlled by a small number of people. By pretending to be a large number of people, adversaries easily could make whatever they want appear popular, and also then get it amplified by a recommender algorithm that was greedy for more engagement.

In a later article, “Fiction is Outperforming Reality,” Paul Lewis at the Guardian wrote, “YouTube was six times more likely to recommend videos that aided Trump than his adversary. YouTube presumably never programmed its algorithm to benefit one candidate over another. But based on this evidence, at least, that is exactly what happened … Many of the videos appeared to have been pushed by networks of Twitter sock puppets and bots.” That is, Trump videos were not actually better to recommend, but manipulation by bad actors using a network of fake and controlled accounts caused the recommender to believe that it should recommend those videos. Ultimately, the metrics they picked, metrics that emphasized immediate engagement rather than the long-term, were at fault.

“YouTube’s recommendation system has probably figured out that edgy and hateful content is engaging.” As sociologist Zeynep Tufekci described it, “This is a bit like an autopilot cafeteria in a school that has figured out children have sweet teeth, and also like fatty and salty foods. So you make a line offering such food, automatically loading the next plate as soon as the bag of chips or candy in front of the young person has been consumed.” If the target of the optimization of the algorithms is engagement, the algorithms will be changed over time to automatically show the most engaging content, whether it contains useful information or full of lies and anger.

The algorithms were “leading people down hateful rabbit holes full of misinformation and lies at scale.” Why? “Because it works to increase the time people spend on the site” watching videos.

Later, YouTube stopped optimizing for watch time, but only years after seeing how much harmful content was recommended by YouTube algorithms. At the time, chasing engagement metrics changed both what people watched on YouTube and what videos got produced for YouTube. As one YouTube creator said, “We learned to fuel it and do whatever it took to please the algorithm.” Whatever metrics the algorithm was optimizing for, they did whatever it takes to please it. Pick the wrong metrics and the wrong things will happen, for customers and for the business.

(This was an excerpt from the draft of my book. Please let me know if you like it and want more.)

Saturday, August 19, 2023

AI letter signers not worried about doomsday AI

An article in Wired, "A Letter Prompted Talk of AI Doomsday. Many Who Signed Weren't Actually AI Doomers":
A significant number of those who signed were, it seems, primarily concerned with ... disinformation ... [or] harmful or biased advice ... [But] their concerns were barely audible amid the furor the letter prompted around doomsday scenarios about AI.
Related, one of the sources for that was a blog post over at Communications of the ACM, "Why They're Worried":
Undesirable model behaviors, whether unintentional or caused by human manipulation ... highly convincing falsehoods that could lead many to believe AI-generated misinformation ... highly susceptible to manipulation ... false content by AI recommendation engines ... can be abused by bad actors.
Despite the hype over AI existential risks from some, these AI experts are worried about uses of AI like flooding the zone with propaganda or making it harder to find reliable information in Google search, practical issues with current deployment of LLMs and ML systems that are getting worse over time.

Challenges using LLMs for startups

Someone (not an AI expert) was asking me about applying large language models (LLMs, like ChatGPT) to a particular product. In case it's useful to others, here's an edited version of what I said.

When thinking about LLMs for a task, I think it's important to consider how LLMs work.

Essentially, LLMs are trying to produce plausible text given the context. It's roughly advanced next word prediction, so given previous words and phrases, and an enormous amount of data on what writing usually looks like around similar words and phrases, predict the next words.

The models are trained by having human judges quickly rate a bunch of output for how plausible it looks. This means the models are optimized for producing output that, on first glance, looks pretty good to people.

The models can imitate writing style, especially for a short period of time, but not particularly well or closely. If you ask LLMs to produce what a specific person would say, including the user of the product, you'll get a plausible-looking but unreliable answer.

The models also have no understanding of what is correct or accurate, so they will produce output that is wrong, potentially in problematic ways.

The models have no reasoning ability either, but they often can produce plausible output to questions that appear to require reasoning by rephrasing memorized answers to similar questions.

Unfortunately, these issues mean you may struggle to get LLMs to produce high quality output for many of the features and products you might be thinking about, even with considerable effort and optimization on the LLMs.

If you're prepared for that, you can try to make customers forgiving of the inevitable errors with a lot of effort on the UI and by limiting expectations, but that's not easy!

Friday, June 30, 2023

Attacking the economics of scams and misinformation

We see so many scams and so much misinformation on the internet because it is profitable.

It's cheap to create bogus accounts. It's cheap to use hordes of accounts to shill your scams and feign popularity. Posting false customer reviews easily can make crap look like it's trustworthy and useful.

Bad actors are even more effective when they manipulate ranking algorithms. When fake crowds of bogus accounts like, share, and click on content, algorithms that use customer behavior -- such as trending, search, and recommenders -- think crap is genuinely popular and show it to even more people.

Today, the FTC announced rules that change the game: "Federal Trade Commission Announces Proposed Rule Banning Fake Reviews and Testimonials."

These rules make it much more risky and costly to create fake reviews of your products. About a third of customer reviews are fake!

These new rules also make it much more costly to manipulate social media using fake accounts and shills. From the FTC: "Businesses would be prohibited from selling false indicators of social media influence, like fake followers or views. The proposed rule also would bar anyone from buying such indicators to misrepresent their importance for a commercial purpose."

Important is this changes the economics of spam for the bad guys. Before, faking and shilling was free advertising for the bad guys and often profitable. Now businesses that shill and use fake followers face risk of much higher costs, likely tipping the balance into making a lot of scams unprofitable.

A paper from a decade ago, "The Economics of Spam", describes how difficult it is already is to make money as an email spammer. Then it summaries interventions like the FTC's recent action brilliantly, saying, "The most promising economic interventions are those that raise the cost of doing business for the spammers, which would cut into their margins and make many campaigns unprofitable."

More risky and less profitable means less of it. This action from the FTC is great news for anyone who uses the internet.

Tuesday, June 13, 2023

Optimizing for the wrong thing

Many companies that think of themselves as data-driven underestimate how easy it is for metrics to go terribly wrong.

Take a simple example. Imagine an executive who will be bonused and promoted if they increase advertising revenue next quarter.

The easiest way for this exec to get their payday is to put a lot more ads in the product. That will increase revenue now, but annoy customers over time, causing a short-term lift in revenue but a long-term decline for the company.

By the time those costs show up, that exec is out the door, on to the next job. Even if they stay at the company, it's hard to prove that the increased ads caused a broad decline in customer growth and satisfaction, so the exec gets away with it.

It's not hard for A/B-tested algorithms to go terribly wrong too. If the algorithms are optimized over time for clicks, engagement, or immediate revenue, they'll eventually favor scams, lots of ads, deceptive ads, and propaganda because those tend to maximize those metrics.

If your goal metrics aren't the actual goals of the company -- which should be long-term customer growth, satisfaction, and retention -- then you easily can make ML algorithms optimize for things that hurt your customers and the company.

Data-driven organizations using A/B testing are great but have serious problems if the measurements aren't well-aligned with the long-term success of the company. Lazily picking how you measure teams is likely to cause high future costs and decline.

Sunday, April 30, 2023

Why did wisdom of the crowds fail?

Wisdom of the crowds summarizes the opinions of many people to produce useful results. Wisdom of the crowds algorithms -- like rankers, recommenders, and trending algorithms -- usefully do this at massive scale.

But several years ago, wisdom of the crowds on the internet started failing. Algorithms started recommending misinformation, scams, and disinformation. What happened?

Let's think about it in more detail. What changed that caused problems for wisdom of the crowds? Why did it change? What can we can we do about it?

Importantly, did anyone find ways to mitigate the problems? If some did fix their algorithms from amplifying misinformation on their platforms, how did they do that? And why didn't everyone fix their wisdom of the crowd algorithms to prevent them from amplifying misinformation?

I have my own answers to these questions, but I'm curious to hear others. If you have thoughts, I'm most curious to hear about whether you think anyone (at least partially) addressed the problems aggravating misinformation on the internet and, if so, why you think others have not.

Netflix and their new streaming with ads

I was wondering how well Netflix's new ad-supported plans are doing. There hasn't been a lot of criticial reporting on it, and I'm sure others are wondering too, so let's go take a look at what we can find.

Their Q1 2023 only has a few details, but it sounds like the $7/month ad plan generates more total revenue than the $10/month basic, but does not appear to be more profitable and does not appear to be getting a lot of subscribers.

It's not surprising that customers aren't in love with the new Netflix ad plan. It's got ads and the catalog is smaller, and it's only $3 more a month to upgrade to no ads in basic.

It's also not surprising that Netflix is able to get at least $3/month in ad revenue from these viewers, though it might be surprising if it was also substantially more profitable given the cost of acquiring and serving those ads.

It'll take more time before we'll know how this goes for Netflix. But so far it doesn't seem like it's much of a success?

Only as good as the data

The Washington Post reports on the data used for ChatGPT and other large language models (LLMs):
We found several media outlets that rank low on NewsGuard’s independent scale for trustworthiness: RT.com No. 65, the Russian state-backed propaganda site; breitbart.com No. 159, a well-known source for far-right news and opinion; and vdare.com No. 993, an anti-immigration site that has been associated with white supremacy.

Chatbots have been shown to confidently share incorrect information ... Untrustworthy training data could lead it to spread bias, propaganda and misinformation.

AI is only as good as its data. Obviously using known propaganda like Russia Today will be a problem for ChatGPT. Generally, including disinformation or misinformation will make the output worse.

AI/ML benefits from thinking hard about high quality data and the metrics you use for evaluation. It's all an optimization process. Optimize for the wrong thing and your product will do the wrong thing.

Monday, April 17, 2023

The biggest threat to Google

Nico Grant at the New York Times writes that Google is furiously adding features to its web search, including personalized search and personalized information recommendations, in an "panic" that "A.I. competitors like the new Bing are quickly becoming the most serious threat to Google’s search business in 25 years."

Now, I've long been a huge fan of personalized search (eg. [1] [2]). I love the idea of recommending information based on what interested you in the past. And I'm glad to see so many interested in AI nowadays. But I don't think this is the most serious threat to Google's search business.

The biggest threat to Google is if their search quality drops to the point that switching to alternatives becomes attractive. That could happen for a few reasons, but misinformation is what I'd focus on right now.

Google seems to have forgotten how they achieved their #1 position in the first place. It wasn't that Google search was smarter. It was that Altavista became useless, flooded with stale pages and spam because of layoffs and management dysfunction, so bad that they couldn't update their index anymore. And then everyone switched to Google as the best alternative.

The biggest threat to Google is their ongoing decline in the usefulness of their search. Too many ads, too much of a focus on recency over quality, and far too much spam, scams, and misinformation. When Google becomes useless to people, they will switch, just like they did with Altavista.

Sunday, April 16, 2023

Ubiquitous fake crowds

The Washington Post writes: "The Russian government has become far more successful at manipulating social media and search engine rankings than previously known, boosting ... [propaganda] with hundreds of thousands of fake online accounts ... detected ... only about 1% of the time."

Fake crowds can fake popularity. It's easy to manipulate trending, rankers, and recommender algorithms. All you have to do is create a thousand sockpuppet accounts and have them like and share all your stuff. Wisdom of the crowds is broken.

This can be fixed, but first you have to see the problem clearly. Then you'll see that you can't just use the behavior from every account anymore for wisdom of the crowd algorithms. You have to use only reliable accounts and toss everything spammy or unknown.

Saturday, March 25, 2023

Are ad-driven business models bad?

There's been a lot of discussion that ad-driven business models are inherently exploitative and anti-consumer. I think that's both wrong and not a helpful way to look at how to fix the problems in the tech industry.

I think the problem with ad-driven models is that it's easy and tempting for executives to use short-term metrics and incentives like clicks or engagement. It's the wrong metric and incentives for teams. But I think the problem is more ignorance, or willful ignorance, of that issue.

In the short-term, for an ad-supported product, ad revenue and profitability does look like ad clicks. In the long-term, ad profitability looks like converting performing ads for advertisers over the lifetime of customers. Those are quite a bit different.

With subscription-driven models, it's more obvious that your metrics should be long-term. With ad-driven models, long-term metrics are harder to maintain, and many execs don't realize they need to. If execs let teams optimize for clicks, they eventually find those clicks have long-term costs as customers start leaving, but unfortunately it's quite costly to reverse the damage once you're far down this path.

In the long-term, I think you can improve the profitability of an ad-driven platform by making the content and ads work better for customers and advertisers (raising ad spend, increasing ad competition for the space, and reducing ad blindness) and by retaining customers longer (along with recruiting new customers). That looks a lot like the strategy for increasing the profitability of a subscription-driven platform. So I don't see much of a difference between ad-supported and subscription-supported business models other than the temptation for executives to inadvertently optimize for the wrong thing.

Saturday, March 18, 2023

NATO on bots, sockpuppets, and shills manipulating social media

NATO has a new report, "Social Media Manipulation 2022/2023: Assessing the Ability of Social Media Companies to Combat Platform Manipulation".
Buying manipulation remains cheap ... The vast majority of the inauthentic engagement remained active across all social media platforms four weeks after purchasing.

[Scammers and foreign operatives are] exploiting flaws in platforms, and pose a structural threat to the integrity of platforms.

The fake engagement gets picked up and amplified by algorithms like trending, search ranking, and recommenders. That's why it is so effective. A thousand sockpuppets engage with something new in the first hour, then the algorithms think it is popular and show crap to more people.

I think there are a few questions to ask about this: Is it possible for social media platforms to stop their amplification of propaganda and scams? If it is possible but some of them don't, why not? Finally, is it in the best interest of the companies in the long-run to allow this manipulation of their platforms?

Saturday, February 25, 2023

Too many metrics and the Otis Redding problem

The "Otis Redding problem" is "holding people, groups, or businesses to too many metrics: They can’t satisfy or even think about all of them at once."

The problem is not just that people don't really know what to do anymore. It's that many people, when faced with this, start doing things that reward themselves: "They end up doing what they want or the one or two things they believe are important or that will bring them rewards (regardless of senior management’s strategic intent)."

That quote is from Stanford Professor Bob Sutton's book Good Boss, Bad Boss, which somehow I hadn't read until recently. I've read all of Bob Sutton's other books too, they're all great reads.

This is just one tidbit from that book. There's lots more in there. On the Otis Redding problem, my read is that Bob's advice is to only pick a 2-3 simple, actionable metrics, but then frequently discuss whether they are achieving what you want and change them if they aren't.

By the way, the name the "Otis Redding problem" comes from the line in his song "Sitting on the Dock of the Bay" where he says, "Can’t do what ten people tell me to do, so I guess I’ll remain the same."

Superhuman AI in the game Go

For a few years now, AI achieved superhuman game playing abilities for Go.

It was quite a milestone for AI. When I was in graduate school, people used to joke that AI for Go was where careers go to die. The game has a massive search space, so had thwarted efforts for decades.

So AlphaGo and similar efforts that beat top-ranked Go players was a very big deal indeed when it happened back in 2016. But now, a amateur-level human player just beat a top-ranked AI at playing Go. He won 14 of 15 games.

Most of the reporting on this has been that the player used an exploit, one hole in the AI strategy, that will easily be closed. But I think this will be harder to fix than most people expect.

AlphaGo and similar techniques work by using deep learning to guide the game tree search, focusing it on moves used by experts. This result says you can't do that, that you need to consider more possible moves.

The human won here by doing moves the AI didn't expect, then exploiting the result. It's not that there is just one hole. It's that doing moves outside of what the AI expects, anything outside of what it has seen in the training data, can result in a bad playing by the AI, which can then be exploited by the human.

Solving that means considering more moves by the opponent, which explodes the game tree search, making the search massively exponential again. I suspect it's going to be hard to fix.

Thursday, February 16, 2023

Huge numbers of fake accounts on Twitter

It seems like this should get more attention, "hundreds of thousands of counterfeit Twitter accounts set up by Russian propaganda and disinformation" that are "still active on social media today."

There has been widespread manipulation of social media, customer reviews, and trending, search ranker, and recommender algorithms using fake crowds.

All of these depend on wisdom of the crowds. They try to use what people do and like to help other people find things. But wisdom of the crowds doesn't work when the crowd isn't real.

Caroline Orr Bueno has some more details, writing that "this is the first we've heard of an ongoing campaign involving such a large number of accounts" and that it is clear this is at "a scale with the potential to mass-manipulate."

Orr Bueno also quotes former Twitter executive Yoel Roth as saying "it's all too cheap and all too easy." This is the core problem with misinformation and disinformation in the last decade.

If it is cheap, easy, and profitable to scam and manipulate using huge crowds of fake accounts, you will get huge numbers of fake accounts. The solution will have to be to make it more expensive, difficult, and unprofitable to scam and manipulate using fake accounts.

Details on personalized learning at Duolingo

There's a new, great, long article on how Duolingo's personalized learning algorithms work, "How Duolingo's AI learns what you need to learn".

An excerpt as a teaser:

When students are given material that’s too difficult, they often get frustrated and quit ... [Too] easy ... doesn’t challenge.

Duolingo uses AI to keep its learners squarely in the zone where they remain engaged but are still learning at the edge of their abilities.

Bloom’s 2-sigma problem ... [found that] average students who were individually tutored performed two standard deviations better than they would have in a classroom. That’s enough to raise a person’s test scores from the 50th percentile to the 98th

When Duolingo was launched in 2012 ... the goal was to make an easy-to-use online language tutor that could approximate that supercharging effect.

We'd like to create adaptive systems that respond to learners based not only on what they know but also on the teaching approaches that work best for them. What types of exercises does a learner really pay attention to? What exercises seem to make concepts click for them?

Great details on how Duolingo maximizes fun and learning while minimizing frustration and abandons, even when those goals are in conflict. Lots more in there, well worth reading.

Massive fake crowds for disinformation campaigns

The Guardian has a good article, "'Aims': the software for hire that can control 30,000 fake online profiles", on fake crowds faking popularity and consensus to manipulate opinion.

Misinformation and disinformation are the biggest problems on the internet right now. And it's never been cheaper and easier to do.

Note how it works. The fake accounts coordinate together to shout down others and create the appearance of agreement. It's like giving one person a megaphone. One person now has thousands of voices shouting in unison, dominating the conversation.

Propaganda is not free speech. One person should have one voice. It shouldn't be possible to buy more voices to add to yours. And algorithms like rankers and recommenders definitely shouldn't treat these as organic popularity and amplify them further.

The article is part of a much larger investigative report combining reporters from The Guardian, Le Monde, Der Spiegel, El Pais, and others. You can read much more starting from this article, "Revealed: the hacking and disinformation team meddling in elections".

Tuesday, January 31, 2023

How can enshittification happen?

Cory Doctorow has a great piece in Wired, "The ‘Enshittification’ of TikTok. Or how, exactly, platforms die." It's about that we regularly see companies make their product worse and worse until it hits a tipping point, then the company loses its customers and starts dying.

Enshittification eventually causes the company to die, so isn't in the best interest of the company. It's definitely not maximizing shareholder value or long-term profits. So why does it happen?

Cory Doctorow does have a bit on the why, but could use a lot more: "An enshittification strategy only succeeds if it is pursued in measured amounts ... For enshittification-addled companies, that balance is hard to strike ... Individual product managers, executives, and activist shareholders all give preference to quick returns at the cost of sustainability, and are in a race to see who can eat their seed-corn first."

That's not very satisfying though. I mean, the company dies. Execs are screwing up. Why does that happen? What can be done about it? That's the question I think needs answering.

Understanding exactly why enshittification happens is important to finding real, viable solutions. Is it purposeful or unintentional on the part of teams and company leaders? Is it inevitable or preventable? If you get the root cause wrong, you'll get the wrong solution.

My view is that enshittification is mostly unintentional. I think it's a result of A/B testing, mistakes in setting up incentives, and teams busily optimizing for what's right in front of them instead of keeping their eye on the prize.

I don't think executives intentionally drive companies into the ground. I think most execs and teams have no idea that this path they are going down will cause such long-term harm to the company. If most really don't want to destroy the company, that leads to different solutions.

Layoffs as a social contagion

Stanford Professor Jeffrey Pfeffer wrote about the recent layoffs at tech companies, saying that it hurts the company in the long-term, but CEOs can't avoid the pressure to join in.
[CEOs] know layoffs are harmful to company well-being, let alone the well-being of employees, and don’t accomplish much, but everybody is doing layoffs and their board is asking why they aren’t doing layoffs also.

The tech industry layoffs are basically an instance of social contagion, in which companies imitate what others are doing. If you look for reasons for why companies do layoffs, the reason is that everybody else is doing it ... Not particularly evidence-based.

Layoffs often do not increase stock prices, in part because layoffs can signal that a company is having difficulty. Layoffs do not increase productivity. Layoffs do not solve what is often the underlying problem, which is often an ineffective strategy ... A bad decision.

For more on the harm, please see my old 2009 post from the last time this happened, "Layoffs and tech layoffs".