File 026521
Cooperating Without Looking: Game Theory Model of Trust and Reciprocal Cooperation (File 026521)
Academic research paper analyzing how individuals decide whether to observe the costs of cooperation and how others' observability of this decision-making process affects trust and reciprocal behavior in repeated interactions.
Summary
This peer-reviewed research paper presents a game theory model called the 'envelope game' to explain why people trust cooperators who act without deliberation more than those who calculate costs. The study demonstrates that an equilibrium exists where cooperators are rewarded for cooperation without looking (CWOL) at temptation thresholds, and shows this equilibrium emerges through subgame perfection and evolutionary dynamics. The authors derive predictions about real-world phenomena including political trust, moral decision-making, and the psychology of commitment.
Cooperating Without LookingMoshe Hoffman, 1,2∗† Erez Yoeli, 2,3∗ Martin Nowak 21 Department of Computer Science and Engineering,University of California at San Diego, La Jolla, CA 920932 Program for Evolutionary Dynamics,Harvard University, Cambridge, MA 021382 Federal Trade Commission,600 Pennsylvania Ave. NW, Washington, DC 20004∗ These authors contributed equally to this work.† To whom correspondence should be addressed; E-mail: hoffman.moshe@gmail.com.Cooperation occurs when we take on costs to help others.A keymechanism by which cooperation is sustained is reciprocity: individualscooperate with those who have cooperated in the past. Inreality, we not only condition on others’ past cooperative actions,but also on the decision making process that leads to cooperation:we trust more those who cooperate without calculating the costsbecause they will cooperate even when those costs are high.Wepropose a game theory model to explain this phenomenon. In ourmodel, player 1 chooses whether or not to cooperate with player 2.Player 1 faces a stochastic temptation to defect and, before choosingwhether to cooperate, also decides whether to “look” at therealized temptation. Player 2 observes not only whether player 1ultimately cooperated but also whether she looked, then decides1whether or not to continue interacting with player 1. We find conditionsin which there is an equilibrium where player 2 chooses tointeract with player 1 only if player 1 cooperates without looking(CWOL) and player 1 chooses to CWOL. We show that this equilibriumis robust to both high degrees of rationality, as modeled bysubgame perfection and learning or evolutionary dynamics, as modeledby the replicator dynamic.Using computer simulations, wealso show that it emerges with high frequency, even in the presenceof other equilibria. Additionally, we show that the ability for player1 to avoid looking, and the ability for player 2 to detect lookingincreases cooperation. We propose this model as a possible explanationfor a number of interesting phenomena, and thereby derivenovel predictions about these phenomena, including why we dislike“flip-flopping” politicians and respect principled people more generally,why people cooperate intuitively, and why people feel disgustwhen considering taboo trade-offs, and why people fall in love.Cooperation occurs when we take on costs to help others. A key mechanism by whichcooperation is sustained is reciprocity: individuals condition their behavior on others’past actions [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14]. In reality, we not only conditionon others’ past actions, but also on the decision-making process that lead to cooperation:we place more trust in cooperators who do not take time to carefully weigh the costs ofcooperation and who do not try to collect data on such costs before deciding whether tocooperate. For example, we are impressed by colleagues who agree to proofread a paperwithout thinking twice, and view with suspicion those who ask, “how long will it take?”before agreeing to attend a practice talk. Such considerations are left out of standard2models of reciprocity, which only attend to cooperative actions and not the deliberationprocess leading up to the action.We develop a simple model to explain why “looking” at the costs of cooperation isviewed with suspicion. The explanation we suggest is quite intuitive: those who cooperatewithout looking (CWOL) can be trusted to cooperate even in times when there aretemptations to defect. While this insight can be captured without the need for a formalmodel, it is less clear that cooperators will choose to not look, since they pay a price bycooperating blindly in tempting circumstances. Moreover, the formal model, as will beseen, helps explicate when CWOL will occur as well as what difference the ability to notlook and to observe others not looking will make.We formalize this idea using what we call the envelope game (see figure 1).Theenvelope game distills an interaction between two individuals, or players, in an uncertainenvironment.Thus, we start by assuming there is a distribution of payoffs with twopossibilities: one in which defection is relatively tempting, and another in which it isnot. The temptation to defect is randomly determined. Defection is not tempting withprobability p and tempting with probability 1 − p. Both players know how likely it is thatdefection is tempting. However, at this point, neither player knows size of the temptation.That is, the temptation is placed inside an envelope and the envelope is sealed withoutthe players knowing its content. Next, we assume that one of the players, player 1,chooses whether to learn the size of the temptation, either via mental deliberation or bygathering information. We model this in a simplified way by assuming that player 1 has adichotomous choice: she can choose to open the envelope and look inside it, or not. If sheopens the envelope and looks, she learns the size of the temptation. If she does not openthe envelope and look, she only knows the distribution of payoffs. Player 1 then chooseswhether to cooperate or defect. We model cooperation as costly to player 1, regardless3of the temptation, but more costly if the temptation is high. Thus, if the temptation islow, player 1 gets a > 0 if she cooperates and c l > a if she defects. If the temptation ishigh, she gets a if she cooperates and c h > c l if she defects. We assume that cooperationbenefits the other player, which we refer to as player 2. We assume, for simplicity, thatplayer 2 gets b > 0 if 1 cooperates and d < 0 if 1 defects, and that these payoffs to player 2are the same regardless of 1’s temptation. For now, we also assume that d < p/(1 − p) · b;that is, defection is sufficiently harmful to player 2 that player 2 prefers to avoid theinteraction if player 1 only cooperates some of the time. Finally, we model player 2’strust in player 1. We model this by giving player 2 the choice of whether to continuethe interaction with player 1. If player 2 continues the interaction, then with probabilityw, all previous steps are repeated, potentially indefinitely. That is, the temptation todefect is randomly chosen again, player 1 chooses whether to look at it, and so on. Withprobability 1 − w, the game ends. The parameter w can be interpreted as the likelihoodof a future interaction or the player’s discounting of future payoffs. Crucially, we assumethat player 2 observes both of player 1’s choices: not only whether 1 cooperated, butalso whether 1 first looked. Our model, therefore, applies to situations where player 1’slooking is somewhat observable, for example, if player 1 can gain additional informationabout the costs of defection by asking player 2 questions, or if player 1 can take time toponder the decision and player 2 can observe player 1’s reaction time. We assume thatplayers maximize their expected payoffs. Note that since we assumed a > 0 and b > 0,both players benefit from a cooperative interaction.In the envelope game, a strategy for player 1 dictates whether she will look and whethershe will cooperate, as a function of whether she has looked and cooperated in the past,and as a function of the temptation if she has just looked. For example, one relativelysimple strategy for player 1 is to look and cooperate whenever the temptation is low.4Similarly, a strategy for player 2 dictates whether he will continue or exit as a functionof everything player 1 has done in the past. One relatively simple strategy for player 2is to continue so long as player 1 has cooperated. If both players play according to thispair of strategies, each time the temptation is low, player 1 cooperates and gets a, player2 gets b, and the game continues with probability w until the first time the temptationis high. Then, player 1 defects and gets c h and player 2 gets d. Some straightforwardcalculations yield the players’ expected payoffs: player 1’s is [ap + c h (1 − p)] / [1 − pw]and player 2’s is [bp + d(1 − p)] / [1 − pw]. Because a strategy specifies players’ moves inevery period and the game’s length is undetermined, infinitely many possible strategiesexist. We are primarily interested in the seven strategies represented in Figure 2, whichshows the expected payoff for each player for each strategy pair of interest.Of particular interest is the strategy pair designated in the top left corner of figure 2),in which player 2 discriminates between cooperators who look and those who do notlook, and player 1 cooperates without looking (CWOL). The following simple argumentdemonstrates that this strategy pair is a Nash Equilibrium (which means no player hasan incentive to unilaterally deviate) whenevera≥ c 1−w lp + c h (1 − p). This condition has anatural interpretation: player 1’s expected temptation from defection is less than the gainsfrom an ongoing cooperative interaction. Thus, player 1 would lose from deviating, forexample by looking, because this would end the lucrative ongoing relationship, reducingplayer 1’s payoff from an expecteda1−wto 0. Player 2 would also lose from deviating, bychoosing to exit, because when player 1 is always cooperating, she expectsrelationship, and exiting yields 0.b1−wfrom theOne potential concern with this equilibrium is that, since player 2 is not worse off bynot attending to looking, he might not do so. This turns out to not be the case. Theintuition is that if there is even a small probability player 1 looks, player 2 is better off5attending to looking. In the appendix, we formalize this intuition by showing that thisequilibrium is subgame perfect, a solution concept used to rule out these kinds of concernsin settings where there is a high degree of rationality [15].We might wonder whether CWOL will emerge in a population of players who are notrational, but are evolving or learning their strategies. There are two reasons to suspectit will not emerge. First, CWOL may be susceptible to the invasion of mutants. Inparticular, the concern is that a player 2 mutant might arise that does not attend towhether player 1 looks, rendering looking irrelevant. This mutant’s fitness is no higher orlower than the incumbent strategy, so may grow by drift. Subsequently, a player 1 mutantwhich looks and defects when the temptation is high would have a fitness advantage andproliferate. While the player 2 mutant would then start dying off since it would now havelower fitness than the incumbent, it is not clear that it would do so fast enough for theplayer 1 mutant to also die off, returning to the CWOL equilibrium. Second, CWOL maybe stable but have such a small basin of attraction that it will never emerge. In particular,there are three other equilibria to consider, which can be seen in figure 2. The first iscomprised of the strategy pair where player 1 always defects, and player 2 always exits(which we refer to as the ALLD equilibrium). It is a Nash Equilibrium for all parametervalues. The second is the strategy pair where player 2 exits if player 1 defects, and player1 cooperates with or without looking (we refer to this as the CWL equilibrium). It is aNash equilibrium whena ≥ c 1−w 2. In fact, there are many more Nash Equilibria, wherethe population mixes between different strategies.Consequently, we employ computer simulations of the replicator dynamic to explorethe evolutionary dynamics of the envelope game. The replicator dynamic is the standardmodel for evolutionary dynamics [16, 17, 18, 19], and also models learning dynamics suchas reinforcement learning or prestige-biased imitation [20]. It describes strategies evolving6over time under the assumption that the rate of reproduction within each populationis proportional to the fitness relative to that type’s other strategies.Since replicatordynamics cannot be solved in closed form, they must analyzed using computer simulationsseeded with randomly chosen strategy frequencies. This method also requires restrictingthe analysis to a few strategies. Therefore, we run our replicator analysis only on therestricted set of strategies represented in figure 2. While this set of strategies is restricted,it includes the potentially destabilizing mutants discussed above.To investigate the frequencies in which different equilibria emerge, we randomly seedthe strategy frequencies many times and record the frequency of each strategy after thepopulation has stabilized. Many mixtures of strategies are behaviorally consistent withCWOL, for example, if a large enough fraction of player 2s exit when player 1s look thenno player 1s will look, even if this fraction is less than 1. Thus, we classify populationfrequencies as being behaviorally equivalent to ALLD, CWL, and CWOL (see appendixfor details). We plot the frequencies of each of these for different values of a in figure 3 (seeappendix for analogous figures for other parameters). As can be seen, CWOL emergesoften when it is a Nash Equilibrium.Next, we identify the conditions under which people will be most likely to avoid lookingand detect looking amongst others. For this, we interpret the conditions under whichCWOL is the only cooperative equilibrium, because, as we show in the appendix, when oneallows for mutations, the CWOL equilibrium no longer emerges in dynamic simulationswhen both CWOL and CWL are equilibria. Recall that CWL is a Nash equilibrium ifand only ifa≥ c 1−w 2. This equilibrium condition has a natural interpretation: in order tosustain CWL, the long term gains to player 1 from the ongoing relationship must sufficefor player 1 to cooperate even when player 1 knows the temptation is high in the currentperiod.Contrast this with the condition that determines if CWOL is an equilibrium:7a≥ c 1−w 1p + c 2 (1 − p). This equilibrium condition has a natural interpretation as well:in order to sustain CWOL, the long term gains to player 1 from the ongoing relationshipmust suffice for player 1 to cooperate when player 1 expects the temptation to sometimesbe high.That is, not looking makes the expected–as opposed to realized–gains fromdefection relevant, in a sense smoothing the temptation to defection. Thus, the rangewhere CWOL is an equilibrium and CWL is not, c 1 p + c 2 (1 − p) ≤a1−w< c 2, has thefollowing interpretation: the expected temptation is low but the maximal temptation ishigh. In the appendix, we confirm this result using our subgame perfections analysis. Wealso use dynamics to show that, CWOL increases relative to CWL when we increase themaximal temptation, but hold the mean temptation constant.We identify a second condition under which people will be most likely to avoid anddetect looking by relaxing the assumption d >p b. Then, in the region where d ≤p1−p1−p b,there is a fourth equilibrium. It is the strategy pair where player 2 always continues ifplayer 1 cooperates when the temptation is low, and player 1 looks and cooperates onlywhen the temptation is low (we refer to this as the ONLYL equilibrium). In contrast,CWOL is an equilibrium for all values of d. CWOL is thus the only cooperative equilibriumin the parameter region d >p b, which has the interpretation: defection is sufficiently1−pharmful to player 2 such that player 2 prefers to avoid the interaction if player 1 onlycooperates some of the time.Note that CWOL is an equilibrium over a wider parameter region than both CWLand ONLYL, and thus that the ability to avoid looking and to detect whether others lookincreases the parameter space over which cooperation is feasible. To see this, consider removingplayer 1’s strategy which consists of not looking. Alternatively, consider removingplayer 2’s strategy where he conditions his behavior on whether player 1 has looked. Ineither case, cooperation is only sustained in a Nash Equilibrium ifa≥ ap + c 1−w 2(1 − p) or8d >p b. These are the same condition needed for CWL and ONLYL, respectively, and,1−pas discussed above, these ranges is strictly smaller than the range of values over whichCWOL is an equilibrium.To summarize, CWOL can occur in equilibrium, and this equilibrium is subgame perfect,is stable, and has a sizable basin of attraction in the replicator dynamics. Moreover,we expect player 2 to prefer interacting with player 1s who do not look and player 1s toactively avoid looking when defection is harmful and the temptation to defect is usuallysmall but sometimes huge. Finally, under these conditions, cooperation can be sustainedonly if player 1 can avoid looking and player 2 can observe whether player 1 looks.While we modeled player 1 looking at the temptation to defect, we could analogouslyhave modeled player 1 looking at the benefits to cooperating. This can be interpreted aslooking to see if anyone is watching, asking what one will get, or calculating the value ofthe ongoing relationship. In this and some similar cases, the modifications to the modelwould be straightforward and analysis would be equivalent.We now apply the model to shed light on a number of interesting phenomena, includingwhy we dislike “flip-flopping” politicians and respect principled people more generally, whypeople cooperate intuitively, why people feel disgust when considering taboo trade-offs,and why people fall in love.We trust candidates for political office whose policies are the result of their convictionsand are consistent over time, and distrust those whose policies are carefully constructedin consultation with their pollsters and who “flip-flop” in response to public opinion, ascaricatured by the infamous 2004 Republican presidential campaign television ad showingJohn Kerry wind-surfing, and tacking from one direction to another. At first glance, thisseems irrational: one would think it a virtue for a politician to flexibly respond to publicopinion once in office. This logic is illustrated by John Maynard Keynes’ famous quote:9“When the facts change, I change my mind. What do you do, sir?” However, if policiesbased on conviction or consistency over time are signals that a candidate does not lookat the benefits of cooperation, this is an indication that the candidate can be trusted insituations where making the decision that is right for his or her constituency comes ata large political cost. Consistent with this, opponents take every opportunity to paintcandidates as flip-floppers who cannot be trusted [21]. This argument generalizes outsideof politics to why we respect people who are “principled” over those who are “strategic”.People intuitively cooperate. That is, when people make decisions rapidly, they aremore likely to cooperate than if they have time to deliberate. Additionally, people whocooperate make quicker decisions than those who defect [22]. The Social Heuristics Hypothesisoffers one explanation for this phenomena: we adopt heuristics to avoid incurringcognitive costs associated with deliberation [23, 24, 25]. In a world with repeated interactions,it is usually worthwhile to cooperate, so individuals may adopt heuristics suchas “always cooperate” or “cooperate as long a situation is not a business interaction.”These same individuals, when serving as laboratory subjects, may apply these heuristicsand cooperate even when it is not worthwhile to do so [22, 26].Our model offers the following alternative explanation for intuitive cooperation. Intuitivecooperation may serve to reduce responsiveness to realized costs. Thus, others shouldtrust intuitive cooperators more than deliberate cooperators. Since intuitive cooperatorsare more trustworthy, this may lead people to evolve or learn to become intuitive cooperators.Note that this explanation suggests that intuitive cooperation may be an optimalresponse to others’ ability to detect deliberation, and not, as a heuristic is, an attemptto avoid cognitive costs. For this explanation to be sensible, it must be the case thatwhether or not a decision is made intuitively or deliberatively is detectable. In fact, it is:deliberative decision making leads to slower reaction time, as well as increased pupil size10and heart rate [27], and sometimes blushing or stammering [28]. In contrast to the SocialHeuristics Hypothesis, our model also predicts that decisions related to cooperation aremore likely to be intuitive than other decisions that are similarly usually worthwhile, andthat intuitive cooperators are trusted more than reflective cooperators. To our knowledge,neither of these predictions has been tested yet.People dislike considering trade-offs related to “sacred values” [29].Sacred valuesare values such as love, liberty, honor, justice, or life, that people treat “as possessingtranscendental significance that precludes comparisons, trade-offs, or indeed any minglingwith secular values” [29].While there is variation in what societies consider sacred,virtually all societies have a concept of sacredness [29].Sacred values are so stronglyimbued in us that we do not find them puzzling prima fascia, yet their existence andorigin remains poorly understood. What makes us treat some values as sacred and whatdifferentiates these values from secular values like free time or money that we more readilytrade off?Our model provides one possible explanation. People who calculate costs of tradingoff against sacred values are less trustworthy when it comes to safeguarding these valuesthan people who consider them sacred and would never calculate the costs of tradingoff against them. Responding with disgust to these “taboo trade-offs” may be one wayto prevent us from interacting with people who make such trade offs and hence are lesstrustworthy, and may also be a way to signal to others that we ourselves would notconsider, and therefore make, such trade-offs.Consistent with CWOL, it is taboo toconsider the trade-off even if one ultimately makes the right choice, and the longer thetrade-off is considered for, the harsher the judgement by observers [29]. As with intuitivecooperation, if people who refuse to consider taboo trade-offs are seen as more trustworthy,this may lead people to evolve or learn to abide by these taboos. If CWOL indeed underlies11the phenomena of taboo trade-offs, then it provides a novel prediction: taboo trade-offswill prevail precisely in situations where there is large but infrequent temptation to defectand defection is harmful, such as selling a child, betraying a country, or sleeping withsomeone for a million dollars. It remains to be shown that taboo trade-offs demonstratethis characteristic. It also provides an important policy prescription regarding policiesforbidding taboo trade-offs, for example, the ban on euthanasia: such policies are sociallysuboptimal, since the benefits of cooperating without looking accrue to the individualswho advocate them, but the costs are borne by society.Finally, our model offers an explanation for emotions such as love, and can thus bethought of as a formalization of Frank (1988). Love has the property that we behavealtruistically towards our partners regardless of what temptations arise [30], as illustratedby the wedding vow, “for better or for worse, for richer, for poorer, in sickness and inhealth.” For example, love causes individuals to ignore other potential mates, even if thosemates are better than one’s current mate, as Shakespeare’s Juliet did when her love forRomeo led her to rebuff the advances of the otherwise-more-suitable Paris.Why does love have this property? Our model suggests that those in love will moreoften be chosen as long-term partners. If this argument in fact underlies love, then itis crucial that love is observable, and that it necessarily cannot be displayed while stillattending to costs. There is evidence consistent with this: related emotions are observable[31], cannot be faked [32], and are relied upon by partners when choosing whether tocooperate [33]. There is also reason to believe love and related emotions would be hardto fake, given their autonomic origins, and the costs of placing their activation underconscious control [30, 28]. However, it remains to be shown that love in particular hasthese attributes, and that it cannot be displayed while attending to costs.Our formalism adds novel insights to Frank’s argument about love. First, our model12clarifies that falling in or out of love depends on the distribution of temptations, but nottheir immediate realizations. This suggests people will fall out of love when there is apermanent change in alternative mating opportunities or relationship costs, but not aone-off temptation. For example, a man may fall out of love with his wife after becomingunexpectedly successful, as many more women will be now be interested in him thanhe anticipated when they first met. Second, the model clarifies that love comes with acost–the cost of ignored temptations–and suggests that this cost must be compensatedwith commensurate investment in the relationship. Only sometimes is it worthwhile forthe recipient of love to compensate a suitor, which explains why people actively avoid thestrong affections of those with whom they do not wish to have long-term relationships.Third, our model clarifies why mere discussions of the costs and benefits of a relationshipor a break-up, for example, suggesting a prenuptial agreement, damage the relationship.Such discussions indicate that one is looking at the costs of the relationship and castdoubt on one’s commitment.These arguments extend to anger. Anger can be thought of as “punishing withoutlooking”. It prevents people from looking at the costs of inflicting harm on others after atransgression, thereby deterring future transgressions.AcknowledgmentsThis research was funded in part by the John Templeton Foundation, Grant RFP-12-11 from the Foundational Questions in Evolutionary Biology Fund, the National ScienceFoundation Grant No. 0905645, and the Army Research Office Grant No. W911NF-11-1-0363. We thank Keri Hu for helpful research assistance. Any opinions expressed in thisarticle are those of the authors and not of the Federal Trade Commission or any individualCommissioner.13References[1] R. L. Trivers, Quarterly review of biology pp. 35–57 (1971).[2] J. W. Friedman, The Review of Economic Studies 38, 1 (1971).[3] R. Axelrod, W. D. Hamilton, Science 211, 1390 (1981).[4] D. Fudenberg, E. Maskin, Econometrica: Journal of the Econometric Society pp.533–554 (1986).[5] D. Fudenberg, E. Maskin, The American Economic Review 80, 274 (1990).[6] K. G. Binmore, L. Samuelson, Journal of economic theory 57, 278 (1992).[7] M. A. Nowak, K. Sigmund, Nature 355, 250 (1992).[8] M. Nowak, K. Sigmund, et al., Nature 364, 56 (1993).[9] R. J. Aumann, L. S. Shapley, Long-term competitiona game-theoretic analysis(Springer, 1994).[10] M. A. Nowak, K. Sigmund, Nature 393, 573 (1998).[11] M. A. Nowak, K. Sigmund, Nature 437, 1291 (2005).[12] M. A. Nowak, science 314, 1560 (2006).[13] H. Ohtsuki, Y. Iwasa, Journal of Theoretical Biology 239, 435 (2006).[14] K. Sigmund, The calculus of selfishness (Princeton University Press, 2010).[15] M. J. Osborne, An introduction to game theory (Oxford University Press, USA, 2003).[16] P. D. Taylor, L. B. Jonker, Mathematical biosciences 40, 145 (1978).14[17] J. W. Weibull, Evolutionary game theory (MIT press, 1997).[18] J. Hofbauer, K. Sigmund, Evolutionary games and population dynamics (CambridgeUniversity Press, 1998).[19] M. A. Nowak, Evolutionary dynamics: exploring the equations of life (Harvard UniversityPress, 2006).[20] D. A. Fudenberg, The theory of learning in games, vol. 2 (MIT press, 1998).[21] J. Markon, Obama fires up crowd in virginia with ‘romnesia’speech, The Washington Post, October 19, 2012, Available: http://articles.washingtonpost.com/2012-10-19/politics/35499645_1_romnesia-obama-fires-economy-with-higher-taxes [Last accessed: August11, 2013].[22] D. G. Rand, J. D. Greene, M. A. Nowak, Nature 489, 427 (2012).[23] H. A. Simon, The quarterly journal of economics 69, 99 (1955).[24] A. Tversky, D. Kahneman, Science 185, 1124 (1974).[25] D. Kahneman, P. Slovic, A. Tversky, Judgment under uncertainty: Heuristics andbiases (Cambridge University Press, 1982).[26] D. Rand, et al., Available at SSRN: h ttp://ssrn. com/abstract 2222683 (2013).[27] D. Kahneman, B. Tursky, D. Shapiro, A. Crider, Journal of Experimental Psychology79, 164 (1969).[28] S. Pinker, NY: Norton (1997).15[29] P. E. Tetlock, Trends in cognitive sciences 7, 320 (2003).[30] R. H. Frank, Passions within reason: The strategic role of the emotions. (WW Norton& Co, 1988).[31] P. Ekman, E. R. Sorenson, W. V. Friesen, Science 164, 86 (1969).[32] P. Ekman, R. J. Davidson, W. V. Friesen, Journal of personality and social psychology58, 342 (1990).[33] L. I. Reed, K. N. Zeglen, K. L. Schmidt, Evolution and Human Behavior 33, 200(2012).16Figure 1: The Envelope GameLowpL1 1 2CC1 - pDLDEHigh(1) (2) (3) (4)17Figure 2: Payoffs for a Restricted Set of Strategies in the Envelope GamePlayer 2Player 1Continue if Player 1Cooperates WithoutLookingContinue if Player 1CooperatesAlways ExitCooperate Without Lookinga1−w , b1−wa1−w , b1−wa, bCooperate With Looking a, ba1−w ,b1−wa, bLook and Cooperate OnlyWhen Temptation is Lowap + c h (1 − p), bp + d(1 − p)ap+c h (1−p)1−pw, bp+d(1−p)1−pwap + c h (1 − p), bp + d(1 − p)Always Defect c l p + c h (1 − p), d c l p + c h (1 − p), d c l p + c h (1 − p), d18Figure 3: Learning Dynamics of the Envelope GameEquilibrium ClassificationPlayer 1 Player 2FrequencyCWLExit if Defect1.00.8All D:CWOLC if Low0.60.4Exit ifLookAlwaysExit0.20.0All D1.00.8CWOL:0.60.40.20.01.00.8CWL:0.60.40.20.0a* a**a19Figure LegendsFigure 1: The Envelope GameWe model non-strategic cooperative behavior using what we call the envelope game. (a)Column 1: The game begins when the temptation to defect is randomly chosen, as indicatedby a notice randomly being placed in the envelope. The temptation to defect is lowwith probability p and high with probability 1 − p. Column 2: Then, player 1 chooseswhether to look (open the envelope) or not. Column 3: Player 1 then chooses whether tocooperate or defect. Player 1 may only condition her action on the realized temptationdetermined in column 1 if she looked. Each time player 1 cooperates, then, regardless ofwhether player 1 looked, player 1 gets a > 0 and player 2 gets b > 0. Each time player1 defects, her payoffs depend on whether defection was tempting. If it was not tempting,player 1 gets c l > a and if it was tempting, player 1 gets c h > c l . In either case, each timeplayer 1 defects, player 2 gets d < 0. Column 4: Player 2, having observed both of player1’s choices, chooses whether to continue or exit. If player 2 continues, with probability w,all previous steps are repeated, potentially indefinitely.Figure 2: Payoffs for a Restricted Set of Strategies in the Envelope GameThis table presents the payoffs for the restricted set of strategies used for the replicatoranalysis. Player 1’s strategies are presented in separate rows, and player 2’s strategiesare presented in columns. The payoffs presented in the intersection of a given row andcolumn are those that the players receive if they play the corresponding strategies. Forexample, consider what happens if player 1 looks and cooperates only if the temptation islow (penultimate row) and player 2 repeats continues when player 1 cooperates (middlecolumn). Then, player 1’s expected payoff is [ap + c h (1 − p)] [1 − pw] (the first entry inthe corresponding cell) and player 2’s is [bp + d(1 − p)] [1 − pw] (the second entry in the20same cell; for details of calculations leading to payoffs, see appendix). This calculationis the result of the following logic: each time the temptation is low, player 1 cooperatesand gets a, player 2 gets b, and the game continues with probability w until the firsttime the temptation is high. We refer to the strategy where player 1 cooperates withoutlooking (top row) as CWOL. We also refer to the strategy pair where player 1 CWOLsand player 2 continues if player 1 CWOLs (first column) as CWOL. We refer to thestrategy pairs where player 1 cooperates with or without looking and player 2 continuesif player 1 cooperates (first and second row, and middle middle column) as CWL. Werefer to the strategy pair where player 1 always defects and player 2 always exits (bottomrow and rightmost column) as ALLD. ALLD is always an equilibrium of the envelopegame. CWOL is an equilibrium if a/(1 − w) > c l p + c h (1 − p). CWL is an equilibrium ifa/(1 − w) > c h . This region is a subset of the region for which CWOL is an equilibrium.Figure 3: Learning Dynamics of the Envelope GameWe apply the replicator dynamic to the envelope game restricted to the strategies representedin figure 2. The replicator dynamic describes strategies evolving over time underthe assumption that the rate of reproduction within each population is proportional tothe fitness relative to that type’s other strategies. The replicator dynamic also modelslearning dynamics such as reinforcement learning or prestige-biased imitation. We run1000 time series with randomly seeded strategy frequencies for a range of values of a, andrecord the frequency with which they stabilize in one of the strategy pairs identified infigure 2, or in a behaviorally equivalent equilibrium, as presented in the simplexes. Wevary the value of a along the x-axis. The y-axis represents frequencies, and each coloredline presents the frequency of the strategy pair. The parameter region where the strategypair is supported as an equilibrium is shaded in light red. CWOL is supported as an21equilibrium in the region a ≥ a ∗ = (1 − w) · [c l p + c h (1 − p)], and emerges often in thisparameter region. CWL is supported as an equilibrium in a strictly smaller parameterregion, a ≥ a ∗ = (1 − w) · c h . Neither CWOL nor CWL emerge in the parameter regionswhere they are not supported in equilibrium.22