Three team models to retire, and the one AI makes worse
Key takeaways
- Tuckman's stages, Belbin's roles and the Hawthorne effect get quoted at every leadership offsite, and each rests on thin, dated or misread evidence. The fourth, the Abilene paradox, holds up, and AI is about to make it worse.
- The team these models describe is changing. In a field experiment with 776 professionals, published in 2026, one person with AI matched a two-person team without it. A five-person team with an agent in it isn't the team Tuckman was describing.
- What replaces the models is team learning run as a weekly discipline, and leaders assessed for the capacity to disagree, including with the AI.
How the models have held up: Tuckman, Belbin, Hawthorne and Abilene.
Tuckman's stages came from a 1965 review of about 50 studies, most of them of therapy and training groups. Tuckman later put the model's popularity down to its catchy stage names, and he had a point: the sequence is memorable, which is its whole power. What later research found is that teams don't move through it in order. Connie Gersick's 1988 study of project teams found long stretches of inertia broken by a sudden shift around the midpoint of the deadline, and a team that has "reached performing" can be back in storming by Thursday after an unannounced reorganisation.
The sequence Tuckman is quoted for, and the pattern Gersick found
Belbin's team roles came out of management games at Henley in the 1970s. The original research found eight roles, and a ninth, Specialist, was added later. The most thorough look at its validity, a 2007 review of 43 studies by Aritzeta, Swailes and Senior in the Journal of Management Studies, found the inventory's convergent validity adequate and its discriminant validity weak: several roles overlap so heavily that the instrument can't reliably tell them apart. It's a decent conversation starter and a poor selection tool, and it gets used as both.
The Hawthorne effect is the most misquoted of the four. The story every manager knows is that output rose at Western Electric's Hawthorne Works, in studies run between 1924 and 1932, because workers knew they were being watched. The data don't carry that story. The workers in the famous relay test room also chose their own colleagues and worked as a small group under a friendlier supervisor. When Richard Franke and James Kaul re-ran the numbers in 1978, they credited most of the gain to managerial discipline, the Depression and the quality of raw materials. When Steven Levitt and John List recovered the original lighting data in 2011, they found the textbook descriptions of its remarkable patterns were fiction. A 2014 systematic review of research-participation effects concluded that "new concepts are needed". Read together, Hawthorne is weak evidence for the power of observation, and the causes of what happened in those rooms are still argued over.
The Abilene paradox, Jerry Harvey's 1974 story of a family driving to Abilene because each member thought the others wanted to, is the exception. Harvey published it as The Abilene Paradox: The Management of Agreement, and that subtitle is the argument. His subject is agreement, and how badly a group can read its own. It's an anecdote with no data behind it, yet as a description of a group agreeing to something no member wants, it holds up better than the three models that came with evidence. It's also about to get worse, for reasons section 04 sets out.
They persist because a new GCC head needs a map, and these are the maps in the room.
None of this is news to organisational psychologists. It's news to many leaders I coach, because the models arrive through leadership programs, and they arrive at the moment a leader is most receptive: the first quarter in a new seat. A site leader running a centre several time zones from the parent, with a team that was assembled before they arrived, wants a way to say where the team is and what comes next. Tuckman gives them a sentence for the parent's quarterly review. Belbin gives them a reason the two architects don't get on. Hawthorne gives them permission to walk the floor.
The models are comforting because they make a complex system look like a sequence. We wrote about what the first quarter actually demands in what GCC heads do in their first 90 days, and very little of it is sequential. The danger is that a leader who believes the team is "in norming" stops looking at the team and starts looking at the model.
How AI changes team dynamics: what the P&G Cybernetic Teammate study found.
The most useful recent team study for a GCC head didn't set out to test any of these models. Fabrizio Dell'Acqua, Ethan Mollick, Karim Lakhani and colleagues ran a pre-registered field experiment in 2024 with 776 Procter & Gamble professionals doing real product-development work. It appeared as NBER working paper 33641 in 2025 and in Organization Science in 2026. An individual with AI did as well as a two-person team without it, and teams with AI were the most likely to produce top 10 percent solutions. The AI groups also spent 12 to 16 percent less time, and the silos between R&D and commercial staff dissolved: each side produced balanced solutions it wouldn't have reached alone.
The frameworks have already moved. The Scrum Guide Expanded, published on 18 January 2026 by Ralph Jocham, John Coleman and Jeff Sutherland, names AI as a non-human stakeholder, lets teams "create agents as AI team members", and keeps one guard rail: "maintain clear human accountability for all outcomes". Scaled Agile's AI-Native SAFe, released in June, describes a cross-functional team of typically three to seven people augmented with AI, and says this "does not imply a reduction in the workforce".
Put those together and the group Tuckman, Belbin and the Hawthorne researchers were describing, people learning to work with each other, is now a smaller group of people learning to work with each other and with something that isn't a colleague, doesn't storm and never needs a role, and the old maps leave it out.
The Abilene paradox now has a passenger that always agrees.
The story is worth telling in full, because the detail is the argument. On a 40-degree afternoon in Coleman, Texas, Harvey is playing dominoes on a shaded porch with his wife and her parents when his father-in-law suggests dinner in Abilene, 53 miles away. Harvey's wife says it sounds like a great idea. Harvey would rather stay where he is, but three people now appear keen, so he says it sounds fine to him and adds that he hopes his mother-in-law wants to go. Of course she does, she says. They drive it in a car with no air conditioning, eat badly in a cafeteria, and get back four hours later, worn out.
Then someone says, untruthfully, that it was a great trip, and the truth comes out one person at a time. The mother-in-law had come only because the other three seemed set on it. Harvey had come to keep the peace. His wife had gone along for the same reason. The father-in-law had never wanted to go at all, and had suggested it because he thought everyone else looked bored. Nobody lied, nobody was overruled and nobody argued. Each person answered a question about what the others wanted instead of saying what they wanted, and every one of them got the answer wrong.
What the family said, and what each of them was thinking
The habit underneath has a name in the conflict literature. Kenneth Thomas and Ralph Kilmann published their conflict-mode instrument that same year, and it puts avoiding in the corner where neither assertiveness nor cooperation is present, the one mode in which nobody's needs are met. Avoiding a disagreement feels like protecting the relationship, which is what makes it easy to keep doing. Writing for Scrum.org, Piyush Rahate sets out what a team looks like once the habit settles in. People stop making their thinking visible, so the team loses the transparency it runs on. Alternatives stop being weighed, and the range of views that justified assembling the team narrows to whatever is easiest to agree with. Agreement arrives without commitment behind it. Morale goes last, because people can tell the difference between a decision they own and one they nodded at.
Five ways to handle a disagreement
The mechanism is private doubt and public assent, and groupthink and confirmation bias both feed on it. Add an AI assistant to the room and there's a new passenger, and by default it's the most agreeable one. The model makers have documented the tendency to tell users what they want to hear, and OpenAI rolled back an update in April 2025 for exactly that reason. A leader who drafts a plan, asks the assistant whether it's sound and gets a warm, articulate yes has driven to Abilene faster.
The Scrum authors saw this coming. Alongside permission to add agents, the expanded guide says AI "can be helpful to deliberately test and challenge the existing thinking". That's the right use, and it only happens when someone gives the assistant the dissent role on purpose: argue against this plan, list what would have to be true for it to fail, tell me who on the team would disagree and why. Left to its defaults, the agent completes the paradox.
For a GCC the risk is concrete. Centres run on consensus with a parent in another time zone, and much of that consensus gets manufactured in the gap between a late-night call and a morning stand-up. An assistant that summarises the call as agreement, because agreement is what the transcript sounds like, becomes the paradox's engine. The fix is the one Harvey proposed: someone has to say, out loud, that they didn't want to go to Abilene.
| Model | What it claimed | What holds up | The practice to run instead |
|---|---|---|---|
| Tuckman's stages | Teams progress through forming, storming, norming, performing | Teams cycle, stall and shift suddenly. The stage names are useful vocabulary | Ask the team where it's disagreeing this week |
| Belbin's roles | Nine roles predict team effectiveness | Useful for conversation, too overlapping for selection | Use it once, as a prompt for what each person thinks they bring, and keep it out of hiring decisions |
| Hawthorne effect | Being observed raises output | The observation story doesn't survive re-analysis, and the real causes are disputed | Let teams shape their own membership where you can, and pick managers for how they treat people |
| Abilene paradox | Groups agree to what no member wants | Holds, and gets worse with an agreeable AI in the room | Assign the dissent role, to a person and to the assistant, before any decision that matters |
What to do instead: team learning as a weekly discipline.
Peter Senge described team learning in 1990 as one of five disciplines, and he chose the word discipline on purpose. The discipline turns on the difference between discussion, where people defend positions, and dialogue, where they surface the assumptions under those positions and examine them together. Decades on, his question through the Center for Systems Awareness is still the practical one: how do we create the conditions for people to work together at their best? Answering it doesn't require knowing which stage the team is in.
A mentor of mine, Lyssa Adkins, whose Coaching Agile Teams trained a generation of coaches, frames leadership as post-heroic: the leader draws on the team's collective intelligence and acts wisely inside complexity instead of pretending to predict it. That's a stance you can practise on a schedule, which you can't do with a model.
In a GCC it looks small and repeatable. Once a week, the people who made one decision look at it again, and the dissent gets named out loud. The retrospective gains one question: what the assistant contributed, where it agreed too easily, and what a human overruled. Managers are chosen and reviewed on how their teams say they are treated, which is the charitable reading of Hawthorne and what 360-degree feedback is good for when it is not turned into a ranking. And mentoring pairs are set across functions, so the silo-crossing the Procter & Gamble study saw happens by design rather than by luck.
The leader worth hiring can name the person who pushed back on them, and what they changed because of it.
Christabel Singh · Chief Marketing Officer and Agile coach · Recruise
Assess the leader for the capacity to disagree, including with the tool.
If the practice is scheduled disagreement, the hire is someone who can hold it. That's assessable, and many senior interviews don't assess it. We've set out in assessing for judgement, not just skill and the reference check is the new skills test how to get at judgement rather than fluency. The same structured interview can carry three more questions. Tell me about a decision you reversed because someone junior disagreed. Tell me about a time the tool or the data was wrong, and how you knew. Who on your last team was the person who said no, and what happened to them? Score the answers on an interview scorecard like anything else. A candidate who can't name a dissenter, or names one and describes managing them out, has told you how the Abilene trips will go.
For leaders reading this from the other side of the table, the same test applies to the seat. If the parent organisation can't name where it disagrees with the centre, or the interviews never surface a hard question, the role may be an Abilene trip already under way. We cover how to read that in is a GCC leadership move right for you, and if you're weighing a move, you can share your profile with our search team.
A closing note on humility. Tuckman, Belbin and the Hawthorne researchers were careful people whose work got flattened into slogans by the people who quoted it. Today's research on AI and teams will be flattened too, and in 10 years someone will write about the AI-era team models to retire. The practice underneath, a team that examines its own thinking and a leader who makes dissent safe and scheduled, hasn't changed in 60 years and won't in the next 10.
Frequently asked questions
Is Tuckman's forming, storming, norming, performing model wrong?
It's unsupported as a sequence. It came from a 1965 review of about 50 studies, mostly of therapy and training groups, and later research, including Connie Gersick's 1988 study of project teams, finds that teams cycle, stall and shift suddenly instead of moving through stages in order. The stage names are a useful vocabulary for what a team is doing this week, and a poor map of what it'll do next.
What is the Abilene paradox?
It is a group agreeing to something no member of it wants. Jerry Harvey named it in a 1974 article, The Abilene Paradox: The Management of Agreement, after his family drove 53 miles across Texas in 40-degree heat for a bad cafeteria dinner, each of the four believing the other three wanted to go. Nobody lied and nobody was overruled. Each person answered what they thought the group wanted rather than saying what they wanted, so the group acted on a consensus that did not exist. It is not groupthink, where a group converges on one view to keep the peace. In Abilene everyone privately disagrees and assumes they are alone in it. In a team it shows up as agreement without commitment, silence read as consent, and decisions that are quietly abandoned later.
How does AI change team dynamics?
The best current evidence is the Cybernetic Teammate field experiment with 776 Procter & Gamble professionals, by Dell'Acqua, Mollick, Lakhani and colleagues, published in Organization Science in 2026. An individual working with AI matched a two-person team without it, teams with AI produced more top-tier solutions, work took 12 to 16 percent less time, and silos between technical and commercial staff dissolved. The risk that comes with it is agreeableness: an assistant that affirms what the group already believes strengthens the Abilene paradox and groupthink unless someone deliberately gives it the job of arguing the other side.
Is the Hawthorne effect real?
Not in the form usually quoted. The popular story says output at Western Electric's Hawthorne Works rose because workers knew they were being watched. Franke and Kaul's 1978 statistical re-analysis credited most of the gains to managerial discipline, the Depression and raw-material quality, and Levitt and List's 2011 analysis of the original lighting data found the textbook patterns were fiction. It's weak evidence for the power of observation, and the causes of the Hawthorne gains are still disputed.
What should a GCC head use instead of team development models?
Team learning in Peter Senge's sense, run weekly. Re-examine one decision with the people who made it and name the dissent out loud. Give the retrospective an agenda item on what the AI assistant contributed and where it agreed too easily. Choose and review managers on how the team says they treat people. And assess senior hires for the capacity to disagree, with people and with the tool, using structured interview questions and reference checks that ask specifically about dissent.
One hiring pattern worth knowing, every ten days.
The Mandate Desk is our read on the senior GCC talent market — one signal that moved, the read behind it, and one thing worth doing. Written from live placement data.
More from Recruise Insights.
GCC Leadership A Director in the US and a Director in your India GCC aren't the same hire
US parents level India roles off their own org chart. The title travels; the span and the decision rights don’t, and the search inherits the difference.
GCC Leadership What Bengaluru GCC heads actually do in their first 90 days
We spoke to GCC heads who have been in role under a year. The published playbook and the actual playbook diverge sharply.
GCC Leadership Hiring authority is the mandate conversation nobody has upfront
The charter says the GCC head owns the centre, yet senior hires still route through headquarters. That gap shows up as slow, compromised hiring.
Have a senior seat to fill?
Tell us the mandate — the role, the level, the market. We’ll come back with what the market is really doing on it, and how we’d run the search.