Most recently I read the book, “the evolution of co-operation”, by Robert Axelrod. It’s a very well informal and practical introduction to the theory of cooperation. Consolidated well to go along with the theory of evolution and explaining cooperation amongst diverse set entities which could make choices, from bacteria to large governments.
Main problem that theory of cooperation tries to address is, ‘under what conditions can cooperation emerge in the world of egoists without central authority? And what will be the effects of the resulting system?’ Centuries ago, a pessimistic view of Thomas Hobbes argued that, “the nature is dominated by selfish individuals who labels life as solitary, poor, short, nasty, brutish, and cooperation cannot develop without a central authority”[1]. However, deeper analysis will show that cooperation can indeed emerge among egoists even without central authorities. And the above deficiencies occur due other intricate reasons, or rather simpler in view of the theory (lesser importance future interactions, as will be discussed below).
The framework is built upon iterated prisoners dilemma, meaning participants can’t go away by making a decision. But also have to consider for future interactions, what ripples their current decision may form. Necessary abstractions for the theory to be defined are as: the payoffs doesn’t have to be comparable; cooperation need not be considered favorable by rest of the world; participants need not be rational and may necessarily not even be taking conscious choices; the participants can communicate with each other only through the sequence of their behavior.
The game of prisoners dilemma involves two participants each with simple choices: to co-operate (C) or to defect (D). Mutual cooperation rewards R to both the participants,
and mutual defection with P, while (cooperation, defection) rewards with (S, T). Here the relation between rewards (payoffs) is such that T>R>P>S and R>(T+S)/2.
(T is the temptation to defect, rewarding the highest, only when other player is not tempted as well). It is well known result in prisoners dilemma that
for any participant the best choice independent of other participant is to defect (Nash Equilibrium). Which will result in lesser reward P for each individual,
but better than reward S for who tries to cooperate. Now, in the iterated prisoners dilemma, the interaction between players could go longer.
Although if the horizon of interaction was defined, the analysis from last interaction to first will again result in an optimal strategy 'series of mutual defections'.
The real world is rarely like that, and horizon isn’t known in advance. Hence, we account for that fact by a probability assigned to the future,
call it discount factor w. Higher the w longer the interaction could go. To confirm, the participants in the dilemma are with the goal to maximize
their cumulative payoff. Call ri the payoff at step i, a function of strategy s. Then at each step of interaction the objective of an individual becomes,
maximizes ri + w * ri+1 + w2 * ri+2 …
= maxs∑i ri+1 * wi
For co-operation to emerge it is important that future co-operations are also highly valued.
Hence, for the evolution of cooperation, it is necessary to w be high enough. That critical value of w depends on the strategy and the parameters T, R, P, S.
e.g. in a case where one participant strategizes permanent retaliation (D) after a defection by its contestant, other participant is better with strategy of
permanent C than a D strategy if,
R/(1-w) > T + wP/(1-w), hence wcritical > (T-R)/(T-P).
Hence it also could be said that, “If w is sufficiently high, there is no best strategy independent of strategy used by the other player.” Since, if w wasn’t so high, player might be better off by defecting and getting away with it.
As an experiment, Robert Axelrod held a computer tournament of this game. Strategies were submitted by participants involving experts from different fields, including game theorists, mathematicians, sociologists, etc. The tournament was held in two rounds with results and insights of first round published before entries of the second round. Amongst highly sophisticated and tedious strategies, the one came out performing the best was a strategy Tit for Tat. In this strategy the participant starts with C and then responds with the same choice as made by other player in the last step. Noticeably, most of the strategies which performed well in this round robin tournament were “nice” strategies, meaning they intended to co-operate and were never the first ones to defect. This was largely because they did well among themselves. Other strategies like one for two (defect only if other player defects twice in a row) were too nice, hence easy to exploit. There were also strategies which involved statistical inference for determining other players strategy and then act accordingly (outcome maximization), this with its stochasticity seemed to overcomplicate behavior and tended to add bias from its initial experience, hence leading bad payoffs. To further understand how may these strategies evolve among future generations, more simulations were run. The rule was best strategies will grow in population, while, least performing will decay. Simulations further concreted previous results that nice strategies were also able to sustain simulated evolution. The very simplicity of these nice strategies is what makes them do better than the rest. They tend to co-operate first, they are provoked instantly to dismay the choice of the other player. These simple rules make them easily identifiable and hence easy to adapt against them.
As nice strategies were getting their foothold with more rounds of evolution, a question to ask is what would happen after everyone starts using the same strategy?
Would their be any incentive for an individual to switch to other strategy? Can any strategy invade such native strategy?
(Here invade means gaining higher payoffs against native strategy than a native strategy would make with another participant using native strategy)
Would native strategy be collectively stable (no strategy could invade it). Collectively stable strategies are important, since they are the only ones
that an entire population can maintain in the long run against any mutant strategy. For high enough w, it could be shown that a strategy like Tit for Tat is
collectively stable, wcritical can be found with analysis similar to above. Well, the bad news is even the strategy all D is also collectively stable,
in fact, for any value of w. Then one may ask, what can we do in such environments? A solution exists in the form of co-operation itself.
So instead of an individual trying a new strategy in population of stable strategy, if a cluster of individuals collectively enter such an environment, they have a chance to
get the cooperation started. Lets analyze this case with an assumption that interactions happen on random. Say, in a native environment, a cluster of individuls with a new strategy
enters, accounting for p proportion of its population. A strategy (A) can invade the native strategy if its expected payoff in population is higher than that of native strategy (B).
i.e.
p V(A|A) + (1 - p) V(A|B) > (1 - p) V(B|B) + p V(B|A)
here V is a function of w. Hence for sufficiently smaller values of w, p any
strategy could be invaded by a cluster. (Here we will not make the asssumption that p is very small, instead analyse for a general p)
An important proposition states that, "If a nice strategy cannot be invaded by a single individual, it cannot be invaded by any cluster of individuals either". This proposition implicitly
assumes p is very small, it can be easily proven with this assumption.
Given: V(A|A) > V(B|A)
Let us assume, V(x|y): payoff by using strategy x against player y
Condition to prove now becomes V(A|A) > p V(B|B) + (1 - p) V(B|A) (1)
now, since A is a nice strategy, V(A|A) >= V(B|B)
{l.h.s. of (1) = V(A|A) = p V(A|A) + (1 - p) V(A|A) >= p V(B|B) + (1 - p) V(A|A) > p V(B|B) + (1 - p) V(B|A)}
hence, cluster of B can only invade A if V(B|A) > V(A|A)
but given: V(A|A) > V(B|A),
hence cluster of B can not invade A
However, the assumption is a bit inconsistent through the book [p63-64 vs p212-213]. Hence for broader proof in evolutionary game theory, we ignore this assumption and
derive for conditions which suffice the correctness of the statement.
V(A|A) = R / (1-w), since A is a nice strategy (hence always cooperates, attaining maximum possible payoff)
we have to find conditions for,
(1 - p) V(A|A) + p V(A|B) > p V(B|B) + (1 - p) V(B|A) (some form of the replicator equation [2], Axelrod p212-213)
r.h.s < p V(B|B) + (1-p) V(A|A):
hence if V(B|B) <= V(A|B), then the proposition is true for any p
But if that’s not the case, then, the proposition would not hold. In general for any strategy A to be robust against new strategy B arriving in a cluster,
(1-p) [V(A|A) - V(B∣A)] > p [V(B∣B) − V(A∣B)]
and this critical value of p is: pcritical < [V(A∣A) − V(B∣A)] / [V(A∣A) − V(B∣A) + V(B∣B) − V(A∣B)]
this analysis is valid for single iteration of evolution only, after some iteration, with change of p, things may change.
Until now, the theory didn’t assume any communication between participants, and co-operation still emerged. There is an interesting case study which demonstates practical implication/application of the idea. It comes from world war I, a co-operation that emerged between either side of the no-man’s-land. This was at the frontlines where British and German battalions were leading the lines in a trench war. Both sides were avoiding mutual punishment, despite efforts from high commands. Such ‘live-and-let-live’ type of behavior seems to be emerged from compassion and beliefs that the enemy was a fellow sufferer. It started with events like extreme weather conditions, where attacks would have led higher damage, or when ration was refilled. It was best for both to avoid conflict which would lead to difficult future. If initiated, such kind of co-operation often continued. Once this kind of behavior was learned in one small sector of the front, it was imitated by units in neighboring sectors, or passed onto battalions taking over the operation. But that wasn’t it, occasionally, one side would launch harmless attacks, to show that the co-operation was not because of weakness, and a defection would have its repercussions. Often two or three for one kind of strategies. Importantly it should be noted that such strategies only emerged when enemies were faced for longer duration and not in other types of wars.
From experiments, theory, and examples it is evident that the shadow of the future (w) plays an important role in emergence of co-operation. And the strategies that do well prefer co-operation, not the first to defect, and provoke instantly if tried to be exploited, so as to provide immediate feedback to the opponent. In fact, these ideas can be found in social structures, such that institutions provide ways to enlarge w so co-operation could emerge. And the other way when co-operation is to be discouraged, as in case of corruptions where co-operation could benefit only participants, and not the social structure as whole. In real life it is also important that the choices made by the other participant to be clearly and quickly recognized. So, strategy could be adjusted to lessen the exploitation of either side. Sometimes the evolution of a strategy may also depend on its territory. If a “nice” strategy was deployed in neighborhood largely surrounding strategy never to co-operate, a “nice” strategy might find itself in comparative lesser payoffs and might shift to nearby strategy receiving higher payoffs.
The evolutionary approach is based on a simple principle: whatever is successful is likely to appear more often in the future. It often relies on tedious trial and error learning, which is slow and painful. But if we understand the process better, we can use our foresight to speed up the evolution. There would not be a single strategy that performs the best, but co-operation helps participants gain higher payoffs.
29 Jun 2025.