# 1. Probability Models and Axioms

https://www.youtube.com/watch?v=j9WZyLZCBzs

[00:00] the following content is provided under
[00:01] a Creative Commons license your support
[00:04] will help MIT open courseware continue
[00:06] to offer highquality educational
[00:08] resources for free to make a donation or
[00:11] view additional materials from hundreds
[00:13] of MIT courses visit MIT opencourseware
[00:17] at
[00:21] ocw.mit.edu
[00:23] okay so uh welcome to 6041 6431 a class
[00:28] on probability models and the like I'm
[00:31] John C I will be teaching this class and
[00:35] I'm looking forward to this being an an
[00:37] enjoyable and also useful experience we
[00:41] have a fair amount of staff involved in
[00:43] this course your restation instructors
[00:46] and also a bunch of T but I want to
[00:48] single out our head ta uzoma who is the
[00:52] key person in this class everything has
[00:55] to go through him if he doesn't know in
[00:57] which recitation section you are then
[01:00] simply you do not exist so keep that in
[01:02] mind all right so we want to jump right
[01:06] in into the subject but I'm going to
[01:08] take just a few minutes to talk about a
[01:10] few administrative details and how the
[01:13] course is run so we're going to have
[01:15] lectures twice a week and I'm going to
[01:18] use oldfashioned transparencies now you
[01:20] get copies of these slides with plenty
[01:22] of space for you to keep notes on them
[01:25] uh a useful way of uh making good use of
[01:28] the slides is to use them as a sort of
[01:32] pneumonic summary of what happened in
[01:34] lecture not everything that I'm going to
[01:37] say is of course on the slides but by
[01:39] looking them you get a sense of what's
[01:41] happening right now and it may be a good
[01:43] idea to review them before you go to
[01:46] recitation so what happens in recitation
[01:48] in recitation your recitation instructor
[01:51] is going to maybe review some of the
[01:53] theory and then solve some problems for
[01:56] you and then you have tutorials where
[01:59] you meet in very small groups together
[02:01] with your ta and what happens in
[02:03] tutorials is that you actually do the
[02:05] problem solving with the help of your ta
[02:08] and the help of your classmates in your
[02:10] in your tutorial section now probability
[02:13] is a tricky subject you may be reading
[02:15] the text listening to lectures
[02:16] everything makes perfect sense and so on
[02:19] but until you actually sit down and try
[02:22] to solve problems you don't quite
[02:23] appreciate the subtleties and the
[02:25] difficulties that are involved so
[02:27] problem solving is a key part of this
[02:29] class and tutorials are extremely useful
[02:32] just for this reason because that's
[02:34] where you actually get the practice of
[02:36] solving problems on your own as opposed
[02:38] to seeing someone else uh who's solving
[02:41] them for you okay uh about mechanics a
[02:45] key part of what's going to happen today
[02:47] is that you will all turn in your uh
[02:50] schedule forms that are at the end of
[02:52] the handout that you have in your hands
[02:55] then the Tas will be working frantically
[02:58] through the night and they're going to
[03:00] be producing a list of uh who goes into
[03:04] what section and when that happens any
[03:08] person in this class with probability
[03:10] 90% is going to be happy with their
[03:12] assignment and with probability 10%
[03:15] they're going to be unhappy now unhappy
[03:18] people have an option though you can
[03:21] resubmit your form together with your
[03:23] full schedule and constraints give it
[03:25] back to the Head ta who will then do
[03:28] some further juggling
[03:30] and reassign people and after that
[03:33] happens 90% of those unhappy people will
[03:36] become happy and 10% of them will be
[03:40] left unhappy okay so what's the
[03:43] probability that out that random person
[03:46] is going to be unhappy at the end of
[03:48] this
[03:48] process it's 1% excellent maybe you
[03:51] don't need this class okay so 1% we have
[03:54] about 100 people in this class so
[03:56] there's going to be about one unhappy
[03:58] person and anywhere you look in life
[04:01] there's all in any group you look at
[04:03] there's always one unhappy person right
[04:05] so what can we do about
[04:08] it all right another important part
[04:10] about mechanics is to read carefully the
[04:12] statement that we have about
[04:14] collaboration academic honesty and all
[04:16] that you encouraged it's a very good
[04:18] idea to work with other students you can
[04:21] consult sources that are out there but
[04:24] when you sit down and write your
[04:26] Solutions you have to do that by setting
[04:28] things aside and just just write them on
[04:31] your own you cannot copy something that
[04:33] somebody else has given to you uh one
[04:37] reason is that we're not going to like
[04:40] it when it happens and another reason is
[04:43] that you're not going to do yourself any
[04:45] favor really the only way to do well in
[04:47] this class is to get a lot of practice
[04:49] by solving problems yourselves so if you
[04:52] don't do that on your own then when quiz
[04:54] and exam time comes uh things are going
[04:57] to be difficult so as I mentioned
[04:59] mentioned here so we're going to have
[05:01] section restation sections that some of
[05:04] them are for 6041 students some are for
[05:06] 631 students The Graduate section of the
[05:09] class now undergraduates can sit in The
[05:13] Graduate recitation sections what's
[05:14] going to happen there is that things may
[05:17] be just a little faster and you may be
[05:19] covering a problem that's a little more
[05:21] advanced and is not covered in the
[05:23] undergrad sections but if you sit in The
[05:25] Graduate section you still just respond
[05:27] and you're an undergraduate you're still
[05:29] just responsible for the undergraduate
[05:32] material that is you can just do the
[05:34] undergraduate work in the class but
[05:36] maybe be exposed at a different
[05:40] section okay um few words about the
[05:44] style of this class we want to focus on
[05:47] basic ideas and Concepts there's going
[05:51] to be lots of formulas but what we try
[05:53] to do in this class is to actually have
[05:55] you understand what those formulas mean
[05:58] and in a year from now when when almost
[06:00] all of the formulas have been wiped out
[06:02] from your memory you still have the
[06:04] basic concepts you can understand them
[06:06] so when you look things up again they
[06:09] will still make sense uh it's not a it's
[06:13] not a plug and chug kind of class where
[06:16] you're given list of formulas you're
[06:18] given numbers and you plug in and you
[06:20] get answers the really hard part is
[06:23] usually to choose which formulas you're
[06:25] going to use you need judgment you need
[06:27] intuition lots of prob ility problems at
[06:30] least the interesting ones often have
[06:33] lots of different solutions some are
[06:34] extremely long some are extremely short
[06:37] the extremely short ones usually involve
[06:39] some kind of deeper understanding where
[06:42] of what's going on so that you can pick
[06:44] a shortcut and use it and hopefully
[06:47] we're going to develop this skill during
[06:49] this
[06:50] class now I I could spend a lot of time
[06:55] in this lecture talking about why the
[06:57] subject is important I'll keep it short
[06:59] because I think it's almost obvious
[07:02] anything that happens in life is
[07:04] uncertain there's uncertainty anywhere
[07:07] so whatever you try to do you need to
[07:09] have some way of dealing or thinking
[07:11] about this uncertainty and the way to do
[07:14] that in a systematic way is by using the
[07:17] models that are given to us by
[07:18] probability Theory so if you're an
[07:20] engineer and you're dealing with a
[07:22] communication system or signal
[07:24] processing basically you're facing a
[07:26] fight against noise noise is random is
[07:29] uncertain how do you model it how do you
[07:31] deal with it if you're a manager yes
[07:35] you're dealing with customer demand
[07:37] which is of course random or you're
[07:38] dealing with the stock market which is
[07:41] definitely random or you play at the
[07:44] casino which is again random and so on
[07:48] and the same goes for pretty much any
[07:49] other field that you can think of but uh
[07:54] independent of which field you're coming
[07:56] from the basic concepts and tools are
[07:59] really all the same so you may see in
[08:01] bookstores that there there are books
[08:04] probability for scientists probability
[08:06] for engineers probability for social
[08:08] scientists probability for astrologists
[08:11] well what all those books have inside
[08:13] them is exactly the same models the same
[08:16] equations the same problems they just
[08:18] make them somewhat different word
[08:20] problems the basic concepts are just one
[08:23] and the same and we'll take this as an
[08:26] excuse for not going too much into to
[08:29] specific domain applications we will
[08:32] have problems and examples that are
[08:34] motivated in some loose sense from Real
[08:36] World situations but we're not really
[08:39] try in this class to develop the skills
[08:42] for a specific uh for domain specific
[08:45] problems rather we're going to try to
[08:47] stick to General understanding of the
[08:51] subject okay so the next slide of which
[08:54] you do which you do have in your handout
[08:57] gives you a few more details about the
[08:59] class
[09:00] uh maybe one thing to comment here is
[09:03] that you do need to read the text in
[09:06] with Calculus books perhaps you can live
[09:09] with just a two-page summary of all the
[09:11] interesting formulas in calculus and you
[09:13] can get by
[09:15] uh just with those formulas but here
[09:18] because we want to develop Concepts and
[09:20] intuition actually reading words as
[09:23] opposed to just browsing through
[09:25] equations does make a difference in the
[09:27] beginning the class is kind of easy when
[09:30] we deal with discrete probability that's
[09:32] the material until our first quiz and
[09:35] some of you may get by without being too
[09:38] systematic about following the material
[09:40] but it does get substantially harder
[09:43] afterwards and I will keep restating
[09:46] that you do have to read the text to
[09:48] really understand the
[09:51] material okay so now we can start with
[09:55] the real part of the lecture let us set
[09:59] the goals for
[10:00] today so probability probability theory
[10:04] is a framework for dealing with
[10:07] uncertainty for dealing with situations
[10:09] in which we have some kind of Randomness
[10:12] so what we want to do is by the end of
[10:14] today's lecture to give you anything
[10:17] that you need to know how to set up what
[10:20] does it take to set up a probabilistic
[10:23] model and uh what are the basic rules of
[10:26] the game for dealing with probabilistic
[10:29] model
[10:30] so by the end of this lecture you will
[10:32] have essentially recovered half of this
[10:34] semester's tuition right so we're going
[10:37] to talk about probabilistic models in
[10:39] more detail the sample space which is
[10:42] basically a description of all the
[10:44] things that may happen during a random
[10:46] experiment and the probability law which
[10:49] describes our beliefs about which
[10:51] outcomes are more likely to occur
[10:53] compared to other outcomes uh
[10:56] probability laws have to obey certain
[10:58] properties that we call the axioms of
[11:00] probability so B main part of today's
[11:03] lecture is to describe those axioms
[11:05] which are the rules of the game and
[11:07] consider a few really trivial
[11:11] examples okay so let's start with our
[11:14] agenda the first piece in a
[11:16] probabilistic model is a description of
[11:18] the sample space of an
[11:20] experiment so we do an
[11:23] experiment and by experiment we just
[11:26] mean that just something happens out
[11:29] there there and that's something that
[11:31] happens it could be flipping a coin or
[11:34] it could be rolling a Dy or it could be
[11:39] doing something in a card game so we fix
[11:42] a particular experiment and we come up
[11:45] with a list of all the possible things
[11:48] that may happen during this experiment
[11:51] so we write down a list of all the
[11:53] possible outcomes so here's a list of
[11:56] all the possible outcomes of the
[11:58] experiment I use the word list but if
[12:01] you want to be a little more formal it's
[12:03] better to think of that list as a set so
[12:07] we have a set that set is our sample
[12:10] space and it's a set whose elements are
[12:13] the possible outcomes of the experiment
[12:15] so for example if you're dealing with
[12:17] flipping a coin your sample space would
[12:19] be Heads This is one outcome tails is
[12:23] one outcome and this set which has two
[12:25] elements is the sample space of the
[12:27] experiment
[12:29] okay what do we need to think about when
[12:32] we're setting up this sample space first
[12:34] the list should be mutually exclusive
[12:36] collectively exhaustive what does that
[12:38] mean collectively exhaustive means that
[12:41] no matter what happens in the experiment
[12:43] you're going to get one of the outcomes
[12:46] inside here so you have not forgotten
[12:49] any of the possibilities of what may
[12:51] happen in the experiment mutually
[12:53] exclusive means that if this happens
[12:56] then that cannot happen so it's at the
[12:59] end of the experiment you should be able
[13:01] to point out to me just one exactly one
[13:05] of these outcomes and say this is the
[13:07] outcome that
[13:09] happened okay so these are sort of basic
[13:12] requirements there's another requirement
[13:15] which is a little more loose when you
[13:16] set up your sample space sometimes you
[13:18] do have some Freedom about the details
[13:20] of the of what you're of how you're
[13:23] going to describe it and the question is
[13:26] how much detail are you going to include
[13:28] so let's let's take this coin flipping
[13:30] experiment and think of the following
[13:32] sample space one possible outcome is
[13:35] heads a second possible outcome is Tails
[13:39] and it's
[13:40] raining and the third possible outcome
[13:43] is Tails and it's not
[13:48] training so this is another possible
[13:51] sample space for the experiment where I
[13:53] flip a coin just
[13:55] once it's a legitimate one these three
[13:58] possib abilities are mutually exclusive
[14:01] and collectively exhaustive which one is
[14:04] the right sample space is it this one or
[14:07] that one well if you think that my coin
[14:10] flipping inside this room is completely
[14:12] unrelated to the weather outside then
[14:15] you're going to stick with this sample
[14:17] space if on the other hand you have some
[14:20] superstitious beliefs that maybe rain
[14:23] has an effect on my coins you might work
[14:27] with a sample space of this kind
[14:29] so you probably wouldn't do that but
[14:32] it's a legitimate option strictly
[14:34] speaking now this example is a little on
[14:36] the bit on the frivolous side but the
[14:39] issue that comes up here is a basic one
[14:41] that shows up anywhere in science and
[14:44] engineering whenever you're dealing with
[14:45] a model or with a situation there are
[14:48] zillions of details in that situation
[14:50] and when you come up with a model you
[14:52] choose some of those details that you
[14:55] keep in your model and some that you say
[14:58] well these are irrelevant or maybe there
[15:00] are small effects I can like neglect
[15:03] them and you keep them outside your
[15:05] model so there's definitely when you go
[15:07] to the real world there's definitely an
[15:09] element of art and some judgment that
[15:12] you need to do in order to set up an
[15:14] appropriate sample
[15:19] space so an easy example
[15:22] now so of course the elementary examples
[15:25] are coins cards dice
[15:29] so let's deal with dyes but to keep the
[15:31] diagram small instead of a six-sided Dy
[15:34] we're going to think about a Dy that
[15:36] only has four faces so you can do that
[15:39] with a tetrahedr and doesn't really
[15:40] matter basically it's a die that when
[15:42] you roll it you get a result which is
[15:45] one two three or four however the
[15:48] experiment that I'm going to think about
[15:50] will consist of two rolls of a
[15:55] Dy crucial Point here I'm rolling the Dy
[15:58] twice
[15:59] but I'm thinking of this as just one
[16:02] experiment not two different experiments
[16:05] not a repetition of twice of the same
[16:09] experiment so it's one big experiment
[16:12] during that big experiment various
[16:13] things will happen such as I'm rolling
[16:16] the D once and then I'm rolling the die
[16:20] twice okay so what's the sample space
[16:23] for that experiment well the sample
[16:26] space consists of the possible outcomes
[16:28] one one possible outcome is that your
[16:31] first role resulted in two and the
[16:34] second role resulted in three in which
[16:37] case the outcome that you get is this
[16:40] one a two followed by three this is one
[16:43] possible
[16:44] outcome the way I'm describing things
[16:48] this outcome is to be distinguished from
[16:50] this outcome here where a three is
[16:54] followed by
[16:56] two if you're playing Bon it does
[16:58] doesn't matter which one of the two
[17:01] happened but if you're do dealing with a
[17:04] probabilistic model that he wants to
[17:06] keep track of everything that happens in
[17:08] this composite
[17:10] experiment there are good reasons for
[17:13] distinguishing between these two
[17:15] outcomes I mean when this happens it's
[17:17] definitely something different from that
[17:19] happening a two followed by a three is
[17:21] different from a three followed by a two
[17:24] so this is the correct sample space for
[17:27] this experiment where we roll the DI
[17:29] twice it has a total of 16 elements and
[17:32] it's of course a finite
[17:35] set sometimes instead of describing
[17:38] sample spaces in terms of lists or sets
[17:41] or diagrams of this kind it's useful to
[17:45] describe the experiment in some
[17:47] sequential way whenever you have an
[17:49] experiment that consists of multiple
[17:51] stages it might be useful at least
[17:54] visually to give a diagram that shows
[17:57] you how those stages evolve and that's
[18:00] what we do by using a sequential
[18:03] description or a tree based description
[18:06] by drawing a tree of the possible
[18:08] Evolutions during our experiment so in
[18:11] this tree I'm thinking of a first stage
[18:13] in which I roll the first die and there
[18:16] are four possible results one two three
[18:19] and four and given what happened in the
[18:22] first in let's say in the first roll
[18:24] suppose I got a one then I'm rolling the
[18:27] second die and the four possibilities
[18:29] for what may happen to the second die
[18:32] and the possible results are 1 two 3 and
[18:34] four again so what's the relation
[18:37] between the two diagrams well for
[18:39] example the outcome two followed by
[18:42] three corresponds to this path on the
[18:45] tree so this path corresponds to two
[18:48] followed by a three any path is
[18:51] associated to a particular outcome any
[18:54] outcome is associated to a particular
[18:56] path and instead of path you may want to
[18:59] think in terms of the leaves of this
[19:01] diagram same thing think of each one of
[19:04] the leaves as being one possible
[19:06] outcome and of course we have 16
[19:09] outcomes here we have 16 outcomes here
[19:12] uh maybe you noticed a subtlety that I
[19:14] used in my language I said I roll the
[19:16] first die and the result that I get is a
[19:19] two I didn't use the word outcome I want
[19:23] to reserve the word outcome to mean the
[19:26] overall outcome at the end of the
[19:29] overall experiment so 2 comma 3 is the
[19:34] outcome of the experiment the experiment
[19:37] consisted of stages two was the result
[19:40] in the first stage three was the result
[19:42] in the second stage you put all those
[19:44] results together and you get your
[19:46] outcome okay perhaps we are splitting
[19:48] hairs here but it's useful to keep the
[19:52] concepts to keep the concepts
[19:55] right what's special about this example
[19:58] is that besides being trivial it has a
[20:01] sample space which is finite there's 16
[20:04] possible total outcomes not every
[20:06] experiment has a finite sample space
[20:09] Here's an experiment in which the sample
[20:11] space is infinite so you are you're
[20:13] playing darts and the target is this
[20:16] square and you're a perfect you're
[20:19] perfect at that game so you're sure that
[20:21] your darts will always fall inside the
[20:25] square so but where exactly your Dart
[20:28] will fall ins inside that square that
[20:29] itself is random we don't know what it's
[20:32] going to be it's uncertain so all the
[20:35] possible points inside the square are
[20:37] possible outcomes of the experiment so a
[20:39] typical outcome of the experiment is
[20:41] going to be a pair of numbers X Y where
[20:44] X and Y are real numbers between zero
[20:47] and one now there's infinitely many real
[20:50] numbers there's infinitely many points
[20:52] in the Square so this is an example in
[20:55] which our sample space is an infinite
[20:58] set
[21:01] okay so we're going to revisit this
[21:04] example a little
[21:05] later okay so these are two examples of
[21:09] what a sample space might be in simple
[21:12] experiments now the more important order
[21:16] of business is now to look at those
[21:18] possible outcomes and make some
[21:20] statements about their relative
[21:22] likelihoods which outcome is more likely
[21:25] to occur to occur compared to the other
[21:29] and the way we do this is by assigning
[21:32] probabilities to the
[21:35] outcomes well not exactly suppose that
[21:39] all you were to do was to assign
[21:41] probabilities to individual outcomes if
[21:44] you go back to this example and you
[21:48] consider one particular outcome let's
[21:50] say this point what would be the
[21:53] probability that you hit exactly this
[21:55] point to infinite precision
[21:58] intuitively that probability would be
[22:00] zero so any individual point in this
[22:03] diagram in any reasonable model should
[22:06] have zero probability so if you just
[22:09] tell me that any individual outcome has
[22:11] zero probability you're not really
[22:13] telling me much to work with for that
[22:17] reason what instead we're going to do is
[22:20] to assign probabilities to subsets of
[22:23] the sample space as opposed to assigning
[22:25] probabilities to individual outcomes
[22:29] so here's the
[22:31] picture we have our sample space which
[22:34] is Omega and we consider some subset of
[22:38] the sample space call it
[22:40] a and I want to assign a num a number a
[22:45] numerical probability to this particular
[22:48] subset which uh represents my belief
[22:52] about How likely this set is to occur
[22:57] okay what do we mean to occur
[22:59] and I'm introducing here a language
[23:01] that's being used in probability Theory
[23:03] when we talk about subsets of the sample
[23:05] space we usually call them events as
[23:08] opposed to subsets and the reason is
[23:11] because it works nicely with the
[23:13] language that describes what's going on
[23:16] so the outcome is a point the outcome is
[23:19] random the outcome may be inside this
[23:23] set in which case we say that event a
[23:27] occurred
[23:29] if we get an outcome inside here or the
[23:31] outcome may fall outside the set in
[23:34] which case we say that event a did not
[23:37] occur so we're going to assign
[23:39] probabilities to
[23:41] events and now how should we do this
[23:45] assignment well probabilities are meant
[23:47] to describe your beliefs about which
[23:49] sets are more likely to occur versus
[23:52] other set so there's many ways that you
[23:54] can assign those probabilities but there
[23:56] are some ground rules for this game
[23:59] first we want probabilities to be
[24:01] numbers between zero and one because
[24:03] that's the usual
[24:05] convention uh so probability of zero
[24:08] means we're certain that something is
[24:09] not going to happen probability of one
[24:11] means that we're essentially certain
[24:13] that something is not going to happen so
[24:15] we want numbers between zero and one we
[24:17] also want a few other things and those
[24:20] few other things are going to be
[24:22] encapsulated in a set of axioms what
[24:25] axioms means in this context it's the
[24:28] ground rules that any legitimate
[24:30] probabilistic model should obey you have
[24:33] a choice of how what kind of
[24:35] probabilities you use but no matter what
[24:38] you use they should obey certain
[24:40] consistency properties because if they
[24:43] obey those properties then you can go
[24:45] ahead and do useful calculations and do
[24:47] some useful reasoning so what are these
[24:50] properties first probabilities should be
[24:53] non-
[24:54] negative okay uh that's our convention
[24:57] we want probability to be numbers
[24:59] between zero and one so they should
[25:00] certainly be non- negative the
[25:02] probability that event a occurs should
[25:04] be a non- negative number what's the
[25:06] second Axiom the probability of the
[25:09] entire sample space is equal to one why
[25:13] does this make sense well the outcome is
[25:17] certain to be an element of the sample
[25:20] space because we set up a sample space
[25:22] which is collectively exhaustive no
[25:24] matter what no matter what the outcome
[25:27] is it's going to be an element of the
[25:28] sample space we're certain that event
[25:31] Omega is going to occur therefore we
[25:34] require we represent this certainty by
[25:36] saying that the probability of Omega is
[25:38] equal to
[25:40] one pretty straightforward so
[25:45] [Applause]
[25:46] far the more interesting axium is the
[25:49] third
[25:51] rule uh before getting into it just a
[25:54] quick reminder if you have two sets a
[25:57] and b
[25:59] the intersection of A and B consists of
[26:02] those elements that belong both to a and
[26:06] to B and we denote it this way when you
[26:09] think probabilistically the way to think
[26:11] of intersection is by using the word
[26:14] end this event this intersection is the
[26:18] event that a occurred and B occurred if
[26:22] I get an outcome inside here a has
[26:24] occurred and B has occurred at the same
[26:27] time so you may find the word end to be
[26:30] a little more convenient than the word
[26:32] intersection and similarly we have some
[26:35] notation for the union of two
[26:38] events which we write this way the union
[26:42] of two sets or two events is the
[26:45] collection of all elements that belong
[26:47] either to the first set or to the second
[26:50] or to both when you talk about events
[26:52] you can use the word or so this is the
[26:55] event that a occurred or or be occurred
[27:00] and this or means that it could also be
[27:02] that both of them
[27:08] occurred okay so now that we have this
[27:10] notation what does the first the third
[27:12] Axiom say the third axum says that if we
[27:17] have two events A and B that have no
[27:20] common
[27:22] elements so here's a here's B and and
[27:28] perhaps this is our big sample space the
[27:31] two events have no common elements so
[27:33] the intersection of the two events is
[27:35] the empty set there's nothing in their
[27:38] intersection then the total probability
[27:40] of a together with B has to be equal to
[27:43] the sum of the individual probabilities
[27:46] so the probability that a occurs or B
[27:49] occurs is equal to the probability that
[27:51] a occurs plus the probability that b
[27:54] occurs so think of probability as being
[27:57] cream cheese
[27:58] you have one pound of cream cheese the
[28:02] total probability assigned to the entire
[28:04] sample space and that cream cheese is
[28:06] spread out over this uh over this set
[28:12] the probability of a is how much cream
[28:15] cheese sits on top of a probability of B
[28:17] is how much sits on top of B the
[28:20] probability of a union B is the total
[28:24] amount of cream cheese sitting on top of
[28:26] this and that
[28:28] which is obviously the sum of how much
[28:30] is sitting here and how much is sitting
[28:32] there so probabilities behave like cream
[28:35] cheese or they behave like mass for
[28:38] example the total mass uh over the mass
[28:43] of a set if you think of some material
[28:47] object the mass of this set consisting
[28:49] of two pieces is obviously the sum of
[28:51] the two masses so this property is a
[28:54] very intuitive one it's a pretty natural
[28:56] one to have
[28:59] okay uh are these axioms enough for what
[29:02] we want to do I mentioned a while ago
[29:05] that we want probabilities to be numbers
[29:08] between zero and one here's an axium
[29:10] that tells you that probabilities are
[29:12] non- negative should we have another
[29:14] Axiom that tells us that probabilities
[29:18] are less than or equal to one it's a
[29:22] desirable property we would like to have
[29:24] it in our
[29:25] hands okay why is it not in that l
[29:29] well the people who are in the axum
[29:31] making business are mathematicians and
[29:33] mathematicians tend to be pretty laconic
[29:36] you don't say something if you don't
[29:38] have to say it and this is a CA this is
[29:41] the case here we don't need that extra
[29:44] Axiom because we can derive it from the
[29:46] existing axioms here's how it
[29:49] goes one is the probability of the
[29:53] entire sample space here we're using the
[29:56] second axom
[30:00] now the sample space is consists of a
[30:05] together with the complement of a okay
[30:08] so this
[30:10] is uh when we I write the complement of
[30:12] a I mean the complement of a inside the
[30:15] set Omega so we have omega here's a
[30:19] here's the complement of a and the
[30:22] overall set is Omega Okay now what's the
[30:27] next step what should I do next which
[30:28] axum should I
[30:30] use we use axium three because a set and
[30:34] the complement of that set are disjoined
[30:36] they don't have any common elements so
[30:39] Axiom 3
[30:40] applies and tells me that this is the
[30:44] probability of a plus the probability of
[30:46] a complement in particular the
[30:50] probability of a is equal to 1 minus the
[30:54] probability of a
[30:56] complement and this is less than or
[30:58] equal to one
[31:02] why because probabilities are non-
[31:05] negative by the first
[31:09] AUM okay so we got the conclusion that
[31:11] we wanted probabilities are always less
[31:14] than or equal to one and this is a
[31:16] simple consequence of the three aums
[31:18] that we have this is a really nice
[31:21] argument because it actually uses each
[31:24] one of those axioms the argument is
[31:27] simple but you you have to use all these
[31:28] three properties to get the conclusion
[31:31] that you want okay so we can get
[31:34] interesting things out of our axioms can
[31:37] we get some more interesting ones how
[31:40] about the union of three
[31:43] sets what kind of probability should it
[31:46] have so here's an event consisting of
[31:49] three of three pieces and I want to say
[31:53] something about the probability of a
[31:56] union B Union C what I would like to say
[32:01] is that this probability is equal to the
[32:03] sum of the three individual
[32:06] probabilities how can I do it I have an
[32:09] axium that tells me that I can do it for
[32:11] two events I don't have an axium for
[32:14] three events well maybe I can massage
[32:16] things and still be able to use that
[32:19] axium and here's the trick the union of
[32:23] three sets you can think of it as
[32:27] forming the union of the first two sets
[32:30] and then taking the union with the third
[32:34] set okay so taking unions you can take
[32:38] the unions in any order that you want so
[32:40] here we have the union of two sets now A
[32:45] B C are
[32:47] disjointed by assumption or that's how I
[32:50] drew it so if a B and C are disjointed
[32:54] then a union B is disjointed from C so
[32:58] here we have the union of two disjoint
[33:00] sets so by the additivity axum the
[33:04] probability of that Union is going to be
[33:06] the probability of the first set plus
[33:09] the probability of the second set and
[33:12] now I can use the additivity axium once
[33:14] more to write that this is probability
[33:17] of a plus probability of B plus
[33:20] probability of
[33:22] C so by using this axum which was stated
[33:25] for two sets we can actually derive
[33:28] a similar property for the union of
[33:30] three disjointed sets and then you can
[33:33] repeat this argument as many times as
[33:35] you want it's valid for the union of 10
[33:37] disjoint sets for the union of 100
[33:40] disjoint sets for the union of any
[33:42] finite number of sets so if A1 up to i n
[33:47] are
[33:50] disjointed then the probability of A1
[33:54] Union a n is equal to the sum of the
[33:58] probabilities of the individual
[34:03] sets Okay special case of this is when
[34:08] we're dealing with finite sets suppose I
[34:11] have just a finite set of outcomes I put
[34:14] them together in a set and I'm
[34:17] interesting in the probability of that
[34:18] set so here's our sample space there's
[34:22] lots of outcomes but I'm taking a few of
[34:25] these and I form a set set out of them
[34:30] this is a set consisting of in this
[34:32] picture of three elements in general it
[34:36] consists of K elements now a finite set
[34:40] I can write it as a union of single
[34:43] element sets so this set here is the
[34:46] union of this one element set together
[34:49] with this one element set together with
[34:51] that one element set so the total
[34:54] probability of this set is going to be
[34:56] the sum of the probabilities of the one
[35:00] element
[35:01] sets now probability of a one element
[35:05] set you need to use the brackets here
[35:09] because probabilities are assigned to
[35:11] sets but this gets kind of tuse so here
[35:14] one abuses notation a little bit and we
[35:17] get rid of those brackets and just write
[35:19] probability of this single individual
[35:23] outcome in any case conclusion from this
[35:26] exercise is that the total probability
[35:30] of a finite collection of possible
[35:33] outcomes the total probability is equal
[35:36] to the sum of the probabilities of
[35:39] individual
[35:41] elements so these are basically the
[35:43] axioms of probability Theory or well
[35:47] they're almost the axioms there are some
[35:50] subtleties that are involved here one
[35:53] subtlety is that this axium here is
[35:57] doesn't doesn't quite do the job for
[35:59] everything we would like to do and we're
[36:01] going to come back to this at the end of
[36:03] the lecture a second subtlety is has to
[36:07] do with weird sets we said that an event
[36:11] is a subset of the sample space and we
[36:13] assign probabilities to
[36:15] events does this mean that we are going
[36:18] to assign probability to every possible
[36:20] subset of the sample space ideally we
[36:24] would wish to do that unfortunately this
[36:27] is not always possible if you take a
[36:30] sample space such as the
[36:33] square the square has nice subsets those
[36:36] that you can describe by cutting it with
[36:38] lines and so on but it does have some
[36:41] very ugly subsets as well that are that
[36:45] are impossible to visualize impossible
[36:47] to imagine but they do exist and those
[36:50] very weird scent sets are such that
[36:52] there's no way to assign probabilities
[36:54] to them in a way that's consistent with
[36:56] the AUM probability okay so this is a
[36:59] very very fine point that you can
[37:02] immediately forget for the rest of this
[37:04] class uh you will only encounter these
[37:07] sets if you end up doing doctoral work
[37:10] on the theoretical aspects of
[37:12] probability Theory so uh so it's just a
[37:16] mathematical subtlety uh that some very
[37:19] weird sets do not have probabilities
[37:21] assigned to them but we're not going to
[37:23] encounter these sets and they do not
[37:25] show up in any applications
[37:29] okay so now let's revisit our examples
[37:32] let's go back to the D example we have
[37:35] our sample space now we need to assign a
[37:38] probability law there's lots of possible
[37:42] probability laws that you can assign I'm
[37:44] picking one here
[37:47] arbitrarily in which I say that every
[37:49] possible outcome has the same
[37:51] probability of 1 over
[37:54] 16 okay why do I make this model well
[37:58] empirically if you have well
[38:00] manufactured dice they tend to behave
[38:03] that way we will be coming back to this
[38:06] kind of story later in this class but
[38:09] I'm not saying that this is the only
[38:12] probability lot that there can be you
[38:14] might have weird di in which certain
[38:16] outcomes are more likely than others but
[38:19] to keep things simple let's take every
[38:21] outcome to have the same probability of
[38:22] one over 16 okay so let's now now that
[38:27] we have in our hands the sample space
[38:29] and the probability law we can actually
[38:31] solve any problem there is we can answer
[38:34] any question that could be posed to us
[38:36] for example what's the probability that
[38:38] the outcome which is this pair is either
[38:41] 1 one or one two we're talking here
[38:45] about this particular event one one or
[38:48] one two so it's an event consisting of
[38:51] these two items according to what we
[38:54] were just discussing the probability of
[38:56] a finite collection of outcomes is the
[38:59] sum of their individual probabilities
[39:01] each one of them has probability one
[39:02] over 16 so the probability of this is 2
[39:05] over
[39:06] 16 how about the probability of the
[39:09] event that X is equal to 1 x is the
[39:12] first rle so that's the probability that
[39:14] the first R uh is equal to one notice
[39:18] the syntax that's being used here uh
[39:22] probabilities are assigned to subsets to
[39:25] sets so we think of this as meaning the
[39:28] set of all outcomes such that X is equal
[39:32] to one how do you answer this question
[39:35] you go back to the picture and you try
[39:36] to visualize or identify this event of
[39:39] Interest X is equal to 1 corresponds to
[39:44] this event here these are all the
[39:46] outcomes at which X is equal to one
[39:48] there's four outcomes each one has
[39:50] probability one over 16 so the answer is
[39:53] 4 over
[39:56] 16 okay
[39:57] how about
[40:00] the probability that x + y is
[40:05] odd okay that will take a little bit
[40:08] more work but you go to the sample space
[40:11] and you identify all the outcomes at
[40:13] which the sum is an odd number so that's
[40:17] a place where the the sum is odd these
[40:20] are other
[40:23] places and I guess that exhausts all the
[40:27] possible outcomes at which we have an
[40:30] odd sum uh we count them how many are
[40:33] there there's a total of eight of them
[40:35] each one has probability 1 over 16 total
[40:37] probability is 8 over 16 and harder
[40:41] question what's the probability that the
[40:42] minimum of the two rols is equal to two
[40:45] this is something that you probably
[40:47] couldn't do in your head without the
[40:50] help of a diagram but once you have a
[40:52] diagram things are simple you ask the
[40:55] question okay this is an event that the
[40:58] minimum of the two rols is equal to two
[41:01] this can happen in several ways what are
[41:03] the several ways that it can happen go
[41:05] to the diagram and try to identify them
[41:07] so the minimum is equal to two if both
[41:10] of them are
[41:12] twos and or it could be that X is two
[41:16] and Y is bigger or Y is two and X is
[41:21] bigger okay I guess we rediscovered that
[41:25] uh yellow and blue make green so we see
[41:29] here that there's a total of five
[41:33] possible outcomes the probability of
[41:35] this event is 5 over
[41:40] 16 simple
[41:44] example but the procedure that we
[41:46] followed in this example actually
[41:48] applies to any probability model you
[41:52] might ever encounter you set up your
[41:55] sample space you make a state M that
[41:57] describes the probability law over that
[41:59] sample space then somebody asks you
[42:01] questions about various events you go to
[42:04] your pictures identify those events pin
[42:07] them down and then start kind of
[42:09] counting and calculating the total
[42:11] probability for those outcomes that you
[42:14] are
[42:15] considering uh this example is a special
[42:18] case of what is called the discrete
[42:20] uniform law the uh model obeys the
[42:23] discrete uniform law if all outcomes are
[42:26] equally likely it doesn't have to be
[42:29] that way that's just one example of a
[42:32] probability law but when things are that
[42:34] way if all outcomes are equally likely
[42:37] and we have n of
[42:41] them and and you have a set a that has
[42:45] little n
[42:48] elements then each one of those elements
[42:50] has probability 1 / capital N since all
[42:54] outcomes are equally likely and for
[42:56] probab to add up to one each one must
[42:59] have this much probability and there's
[43:01] little n elements that gives you the
[43:03] probability of the event of Interest so
[43:06] problems like the one in the previous
[43:08] slid and more generally of the type
[43:10] described here under a discrete uniform
[43:12] law these problems reduce to just
[43:14] counting how many elements are there in
[43:16] my sample space how many elements are
[43:18] there inside the event of Interest
[43:21] counting is generally simple but for
[43:23] some problems it gets pretty complicated
[43:26] and uh you a couple of weeks we're going
[43:28] to have to spend the whole lecture just
[43:30] on the subject of how to count
[43:32] systematically now the procedure we
[43:34] followed in the previous example is the
[43:37] same as the procedure you would follow
[43:39] in continuous probability problems so
[43:41] going back to our Dart problem we get a
[43:43] random Point inside this Square that's
[43:46] our sample space we need to assign a
[43:49] probability law for lack of imagination
[43:51] I'm taking the probability law to be the
[43:54] area of a subset so if we have two
[43:58] subsets of the sample space that have
[44:01] equal areas then I'm postulating that
[44:04] they are equally likely to occur the
[44:06] probability that I fall here is the same
[44:08] as the probability that I fall
[44:10] there the model doesn't have to be that
[44:13] way but if I have sort of complete
[44:15] ignorance of which points are more
[44:16] likely than others that might be the
[44:19] reasonable model to use so equal areas
[44:22] mean equal probabilities uh if the area
[44:25] is twice as large the probabilties going
[44:27] to be twice as big so this is our
[44:31] model uh we can now answer questions
[44:34] let's answer the easy one what's the
[44:36] probability that the outcome is exactly
[44:39] this point that of course is zero
[44:43] because a single point has zero area and
[44:47] since probability is equal to area
[44:49] that's zero probability how about the
[44:52] probability that the sum of the
[44:54] coordinates of the point that we got got
[44:57] is less than or equal to 1/2 how do you
[45:00] deal with it well you look at the
[45:02] picture again at your sample space and
[45:05] try to describe the event that you're
[45:06] talking about the sum being less than
[45:09] 1/2 corresponds to getting an outcome
[45:12] that's below this line where this line
[45:15] is the line where X + Y = to 12 so the
[45:20] intercepts of that line with the axis
[45:22] are 1/2 and
[45:24] 1/2 so you describe the event visually
[45:28] and then you use your probability law
[45:30] the probability law that we have is that
[45:32] the probability of a set is equal to the
[45:35] area of that set so all we need to find
[45:37] is the area of this triangle which is 12
[45:40] * 1 12 *
[45:44] 1/2 equals to 1/
[45:48] 18 okay moral from these two examples is
[45:51] that it's always useful to have a
[45:52] picture and work with a picture to
[45:55] visualize the event that you're talking
[45:57] about and once you have a probability
[46:00] law in your hands then it's a matter of
[46:02] calculation to find the probabilities of
[46:05] an event of Interest the calculations we
[46:07] did in these two examples of course were
[46:09] very simple sometimes calculations may
[46:11] be a lot harder but it's a different
[46:15] business it's a business of calculus for
[46:17] example or being good in algebra and so
[46:20] on as far as probability is concerned uh
[46:23] it's clear what you will be doing and
[46:25] then maybe you're faced with a harder
[46:26] algebraic part to actually carry out the
[46:29] calculations the area of a triangle is
[46:32] easy to compute if I had put down a very
[46:34] complicated shape then you might need to
[46:37] solve a hard integration problem to find
[46:39] the area of that shape but that's stuff
[46:41] that belongs to another class that you
[46:43] have presumably Mastered by now good
[46:46] okay so now let me spend just a couple
[46:48] of minutes to return to a point that I
[46:50] raised before I was saying that the axum
[46:53] that we had about
[46:55] additivity might not not quite be enough
[46:58] let's illustrate what I mean by the
[47:00] following example think of the
[47:02] experiment where you keep flipping a
[47:04] coin and you wait until you obtain heads
[47:06] for the first time what's the sample
[47:08] space of this experiment you might have
[47:11] it might happen in the first flip it
[47:13] might happen in the 10th flip heads for
[47:15] the first time might occur in the
[47:17] millionth flip so the outcome of this
[47:19] experiment is going to be an integer and
[47:21] there's no bound to that integer you
[47:23] might have to wait very much until that
[47:26] happens so the natural sample space is
[47:28] the set of all possible integers
[47:31] somebody tells you uh some information
[47:34] about the probability law the
[47:36] probability that you have to wait for n
[47:38] flips is equal 2 to the minus n where
[47:41] did this come that's a separate story
[47:44] where did it come from somebody tells
[47:46] this to us and they ask and those
[47:49] probabilities are plotted here as a
[47:51] function of N and you're asked to find
[47:53] the probability that the outcome is an
[47:54] even number how do you bow how do you go
[47:57] about calculating that probability so
[48:00] the probability of being an even number
[48:01] is the probability of the subset that
[48:04] consists of just the even numbers so it
[48:08] would be a subset of this kind that
[48:11] includes two four and so one so any
[48:14] reasonable person would say well the
[48:17] probability of obtaining an outcome
[48:19] that's either two or four or six and so
[48:22] on is equal to the probability of
[48:24] obtaining a two plus the probability of
[48:26] obtaining a four plus the probability of
[48:28] obtaining a six and so on these
[48:31] probabilities are given to us so here I
[48:34] have to do my algebra I add this
[48:36] geometric series and I get an answer of
[48:39] 1/3 that's what any reasonable person
[48:42] would do but a person who only knows the
[48:45] axioms that I posted just a little
[48:49] earlier may get stuck they would get
[48:52] stuck at this point how do how do we
[48:54] justify this
[48:59] we had this property for the union of
[49:01] disjointed sets and the corresponding
[49:04] property that tells us that the total
[49:07] probability of finitely many things
[49:10] outcomes is the sum of their individual
[49:12] probabilities but here we're using it on
[49:16] an infinite collection the probability
[49:18] of infinitely many points is equal to
[49:22] the prob to the sum of the probabilities
[49:24] of each one of these to justify by this
[49:27] step we need to introduce one additional
[49:30] rule an additional axum that tells us
[49:32] that this step is actually legitimate
[49:36] and this is the countable additivity
[49:38] axium which is a little stronger or
[49:41] quite a bit stronger than the additivity
[49:43] axium we had before it tells us that if
[49:46] we have a sequence of sets that are
[49:49] disjoined and we want to find their
[49:51] total
[49:52] probability then we are allowed to add
[49:56] their indiv visual probabilities so the
[49:58] picture might be as follows we have a
[50:01] sequence of sets A1 A2 A3 and so on I
[50:07] guess in order to fit them inside the
[50:09] sample space the sets need to get
[50:10] smaller and smaller
[50:12] perhaps uh they are disjointed we have a
[50:15] sequence of such sets the total
[50:17] probability of falling anywhere inside
[50:20] one of those sets is the sum of their
[50:23] individual
[50:25] probabilities key s that's involved here
[50:28] is that we're talking about a sequence
[50:31] of events if by sequence we mean that
[50:35] these events can be arranged in order I
[50:38] can tell you the first event the second
[50:41] event the third event and so on so if
[50:43] you have such a collection of events
[50:45] that can be ordered as first second
[50:48] third and so on then you can add their
[50:51] probabilities you can find to find the
[50:54] probability of their Union so this point
[50:56] is actually actually a little more
[50:57] subtle that you might appreciate at this
[50:59] point and I'm going to return to it at
[51:01] the beginning of the next lecture for
[51:04] now enjoy the first week of classes and
[51:07] have a good weekend thank you
