# Lecture 01 : Introduction to Statistical and Machine Learning

https://www.youtube.com/watch?v=2OXUQp57CMw

[00:23] Hello everyone.
[00:25] Uh, very warm welcome to this, uh, very first lecture of the course, uh, statistical and machine learning methods, basics and applications.
[00:34] So we are at very first module, module one, and in this module, uh, we will discuss about the role of statistics, machine learning, and data science.
[00:45] So these three are interlin, of course, and in this module we'll try to understand, uh, through discussion what are their, uh, kind of, uh, linkage or association and how these three can be utilized in a beneficial, uh, manner.
[01:04] Uh, so in this lecture we are, uh, specifically focusing on the introduction to the statistical and, uh, machine learning.
[01:12] So these two terms we are picking up for this very first lecture.
[01:18] So our focus for this, uh, particular, uh, lecture will be to highlight the complimentary role for
[01:24] This statistical and machine learning.
[01:26] So that means this uh statistical word and this machine learning uh uh this machine word can be used in a combination of the statistical learning as well as this machine learning.
[01:39] So this is we will try to understand.
[01:42] So at the very starting I have uh written here that is statistical and or machine learning.
[01:50] So there are uh two uh major verticals we can just imagine now one is the statistical learning and other one is the machine learning and of course we can utilize these two things in their combination also.
[02:06] So what is this all about?
[02:08] So if I consider them either separately or combinedly then we we should at the very starting we should know that these are all about the methods or algorithms that extracts valuable information from the data.
[02:23] So here I want to focus on the point that
[02:26] Is the information that is available from the data and this extraction of this valuable information is the overall focus for this entire uh effort that we want to put.
[02:40] So the very first question that comes to the students mind that is this automatic whenever we are talking about this machine learning then we need some input and some output will come.
[02:51] So is this uh process becomes automatic?
[02:53] The answer is straightforward.
[02:55] No.
[02:55] And I want to underline this.
[02:59] So we need to understand the general purpose of it.
[03:02] The general purpose of any statistical or machine learning models in such a way that these models should be applicable to the data sets apart from the data sets that was used to train the model.
[03:17] Maybe some of the words at the very starting may sound very new.
[03:19] For example the training so that any model that we develop generally in the machine learning domain we will
[03:28] See in the subsequent lecture the training testing validation similarly for the statistical models we need to develop the model first.
[03:37] So while we develop this kind of model we use some data set so that model knows that whatever the data is used to develop that model to train that uh model but as I am discussing this general purpose.
[03:52] So this general purpose is not only to capture the relationship that is already supplied to the uh model while developing it but it should be able to generalize in capturing the relationship for some new data sets also.
[04:11] That's why it is written that the data set apart from the data set that was used to train or develop the model.
[04:21] So there are uh three verticals or three interlin aspects for all these uh approach when we are whenever we are.
[04:29] Talking about either statistical learning or machine learning or a combination thereof.
[04:38] The first is data, second one uh is model and third one is learning.
[04:46] So when we talk about this uh data it is the core of this any learning as you know that uh so it learns from the data itself.
[04:54] Sometimes these are called the datadriven models also when you are thinking about a model the models are uh related to the process that generates the data.
[05:07] So we have the data whether it is simulated or we have collected from the real observation.
[05:13] So that data is driven by some unknown or or semi-nown processes.
[05:22] So basically whenever we are talking about this modeling or the models that models should try to understand or try to
[05:29] Develop that underlying process and when we call about this learning, this learning is the process of optimizing the parameters of this model.
[05:41] So here, any model that we conceptualize or we try to develop, there are some uh parameters are there, those are those are estimated through different processes of course, that information is coming from the data.
[05:59] So these three interlin aspects when you are trying to connect each other.
[06:06] So when we develop a model, then is model is develop is said that it has learned from the data.
[06:12] This learn word is very important here.
[06:14] It we call that this model has learned from the data if its performance on a given task.
[06:21] So here the task may be simulation or inference or prediction, many things on a given task it improves when the data is taken into account and a good.
[06:33] Model the good is again the subjective word.
[06:37] The good model generalize well to the unseen data.
[06:44] So that means the data that was not given to the model earlier it should also be able to uh capture or do the or or yield the similar performance for the new data sets as well.
[06:59] Now if we have a deeper look about these three interlin aspects.
[07:05] The first one as I mentioned it is the data and as uh it is already already known that is the most precious uh information that uh is uh that is considered in this particular domain of uh study.
[07:22] And generally whenever we are going for some kind of model development and all we consider this data in a vector form whether it can be a uh one-dimensional or multi-dimensional but generally we
[07:35] Consider this data set as a vector.
[07:38] Second one is model again.
[07:42] So as I mentioned, the model is supposed to describe or represent the underlying process, the underlying process to capture.
[07:55] So whether it can capture that process or not, and of course it is learning that process through the data that has been given to this model to learn.
[08:04] Uh, it is used for generating the data, but at the same time it should preserve the information that is available in the available data set.
[08:17] Uh, so this preserving the information available that can have different sense.
[08:21] One is that it should preserve that its statistical properties.
[08:25] It should preserve the underlying uh the the process through which it has been uh develop or it can, it should also.
[08:36] Preserve if there is any u the the correlation or association or causal effect is already there among different variables.
[08:47] The third of these three uh aspects there is the learning and this learning is all about estimating the model parameters and this model parameters can be both it can be static or it can be dynamic using the available data again I repeat so these things initially done with this available data with respect to one utility function I'll come to this utility function in a minute that utility function that evaluates that how well the model capture the underlying process.
[09:24] So how well it is capturing that is not only judged by the data that is already available already given to the model but it also fits well with the
[09:37] Unseen data.
[09:39] So once again very quickly if we see these three aspects, first one is a data which is a most precious commodity in this domain of learning, and we generally treat this data as a vector.
[09:52] Second one is a model.
[09:55] This model is supposed to describe or represent the underlying process to capture it, is used to used for the generating the data preserving the information available in the available data set and learning.
[10:06] Learning is all about the estimating the parameters of this model, and of course I can put it as one s here that is a parameters there, of course there are more than one parameter, but the parameters can be static and it can be dynamic also depending on the what type of modeling approaches we are we are following.
[10:26] So this is all about estimating the parameter using the available data through one utility function which I'm coming in a minute that evaluates that how well the model captures underlying process and it
[10:39] Should fit well for the training as well as for the unseen data.
[10:47] Now coming to that utility function, if I want to discuss in a general perspective, let us say that our response variable is y and there are p different predictor variables are there, and these are denoted as x1, x2, up to xp.
[11:09] So any modeling approach, whether it is statistical or machine learning approach, our target always is to estimate one function that is not known to us, is that function that operates on all those input variables to estimate the response y, and of course there will be some error because none of these utility functions will perfectly fit.
[11:37] So there will be some error term will also be there.
[11:35] So this f is fixed.
[11:42] Of course if we can understand that exact process that is uh there in place, then it is a fixed but it is an unknown function to us which I was referring to that uh uh utility function and it represents the systematic information that the x provides about this x.
[12:02] So that means through this transformation of this f, I can utilize the information that is available to us which I'm calling as this predictors and uh that will help me to estimate about the response here.
[12:19] This epsylon as I mentioned is the random error term which is when which must be independent of the x that is the predictors and it is having a mean zero.
[12:32] The function f that I am mentioning it helps in multiple uh purposes that I mentioned few few minutes back: is it can be used for this prediction, it can be used for the simulation, or it can be used for the
[12:43] Inferences and all these applications we'll see, but from this slide the take-home is that so this f is fixed.
[12:51] But this is not known to us now, so any modeling approach tries to estimate what is the form of this particular uh function f.
[13:04] Now suppose that for a modeling aspect I estimate one function, and I'm just putting I'm writing in that in terms of giving this hat on this f.
[13:12] That means so it is not the exact one, but I have somehow estimated that function.
[13:20] So here what we can uh just for this uh demonstration sake, you can refer to this right side diagram where I'm considering two inputs x1 and x2 because only in three-dimensional the things can be uh shown.
[13:34] So and this is your that response variable, and these red dots in this diagram those are the data points.
[13:44] If we want to fit a function so that means this fcap fcap means that is the estimate.
[13:49] So this yolo surface that you can see on the on this right hand side diagram this yolo surface is basically is given by this a cap.
[14:00] So this is one of the estimate of this underlying uh function just for the demonstration purpose and of course there are some error which are shown in this that vertical uh lines from the surface is basically the distance along the uh response variable that is y from this uh surface to that data point.
[14:20] So those are the error terms here.
[14:25] Now if we just take this uh function as that expectation of this y.
[14:31] So y you remember that last slide we discussed is that actual observation whereas this ycap is the estimate that we get if we assume that f_cap is the uh right function then this estimates of this ycap I'm just
[14:45] Putting so this y minus ycap is giving the error.
[14:48] And that error I'm making it square and taking their expectation.
[14:54] This e stands for the statistical expectation.
[14:57] Uh, and if we just place this y.
[15:01] So what we just started with is that f_sub_x plus epsylon that was shown in this last slide, if you remember.
[15:08] So this is the one that was our initial uh we started with this conceptualization.
[15:14] And then this a prime cap is the expectation of this ycap.
[15:19] Because as I told that this one is having a mean zero, that means if I take the expectation of this epsylon it will be zero.
[15:30] So here uh uh so this uh expectation of this ycap we are getting this that f_cap of x.
[15:36] Slight rearrangement we can do from here.
[15:38] So this fx I'll take and fcap I'll take here in this part that square.
[15:44] And this expectation of this error term is
[15:46] Basically leading to this variance of this error.
[15:51] Now what happens this total error that we started with we have separated into two parts.
[15:56] So this first part is known as the reducible error and the second part is known as that irreducible error.
[16:04] That means if I change this fcap of course these distances these vertical distances here in this right hand side diagram as I have shown those errors that can be reduced.
[16:15] That means the better the surface of estimate of course this uh this error should this part of this error will reduce but this side it will not be reduces.
[16:26] That means so our focus for any of these modeling approaches is to estimate this f with the aim that we can minimize this reducible error part.
[16:36] So that is the overall background philosophy and it is applicable for the both the any statistical modeling scheme or any the machine learning scheme.
[16:44] We'll see that where all there's these.
[16:47] Different things can be applied.
[16:49] So this discussion is just as a general discussion for the overall domain.
[16:55] So before we proceed or dive into this more and more, this both these deeper part, there are some questions that should be interesting to know though.
[17:07] So the first question comes that whenever we are talking about the predictors, then what are the predictors?
[17:14] So some some examples, some real life examples applications we can take.
[17:19] So for that particular variable that we are targeting, so here we are calling as a corresponding response, then what predictors that we utilize.
[17:30] So both in the statistical and this machine learning, this is a big question because of the fact that we'll in a minute we'll show you that how the dimension changes so far as the parameters are concerned and the sample sizes are concerned.
[17:45] So if I just give you an say for example to start with, if I just
[17:47] Give an example that uh heat waves for example.
[17:51] So what are the predictors that we should we should uh select first?
[17:53] Is there any relationship between those things?
[17:59] So this relationship how to ascertain?
[18:01] So either that can be conceptualized or that can have some physical justification, those things we have to see.
[18:09] Now even if there are some relationship, is the relationship linear?
[18:13] So there are n numbers of question: is the relationship linear or nonlinear, or is it static or is it dynamic, is it changing over the the time, is it time varying or not?
[18:25] Those questions will come.
[18:28] So these questions will basically help to conceptualize what type of that function f that we we discuss we should take.
[18:34] So that means that what type of model for a particular uh uh problem given in hand we should uh try.
[18:43] So what are the underlying processes?
[18:45] So if we know some of these underlying processes, definitely as you can understand uh uh
[18:49] This selection uh will be easier.
[18:52] What is the data?
[18:55] So sometimes some of the uh processes data is straightforward but some of the even more uh particularly the the natural problems sometimes that which data what quality of this data what length of this data those things are becoming important.
[19:13] Second that is this data sufficient for the for the inference.
[19:17] So what should be the proper length of this data for a particular modeling aspect.
[19:22] So as I told that these questions are many questions may come that's why I just put a kind of that is not that fully exhaustive list of these questions that may come to our mind when we are just starting but one specific point that we should remember at the starting it is that uh the meaningful information I repeat once again the information is there but the meaningful interpretable that meaningful information uh include the information which can be as simple as the basic.
[19:50] Statistics.
[19:53] Maybe I can say that the climatology of a region or it can be as complex as when we try to draw some inference based on the whatever information it is always the partial information that is available to us or some decision we have to take on a hypothesis.
[20:12] So this information when you just talk about so this information needs to be the meaningful information to us.
[20:21] Maybe these are some of the general discussion that we are having before.
[20:24] So we should have this kind of things clearer before we jump in on some of these application problems.
[20:33] So we are discussing both the statistical learning as well as this machine learning approaches.
[20:41] There are different statistical methods are also there maybe in a combined manner so far but there are some differences are also there between these.
[20:52] two. So it's most important uh the
[20:56] differences it's basically the how we
[20:58] treat the predictor variables in both
[21:01] the domains. This predictor selection
[21:03] when we are talking about that is we
[21:06] choose a the relevant predictor
[21:08] variables from a pool of predictors for
[21:11] example. So sometimes we don't have the
[21:13] entire uh background knowledge so that
[21:16] we have the possible pool but from there
[21:18] we want to select uh some of the most
[21:21] effective predictors. So that is
[21:24] commonly practiced in the statistical uh
[21:26] approaches. So okay fine let mean rather
[21:30] than taking the entire pool of course it
[21:32] is increasing the dimensions and all but
[21:34] sometimes it is avoided in this machine
[21:36] learning. There are some reasons of
[21:38] course because it can handle uh mean
[21:42] much higher dimension as we can see from
[21:44] this from the from the statistical
[21:46] approaches or learning till few I should
[21:49] say 10 years before we generally say
[21:51] that before we start with any
[21:52] statistical model we have to see whether
[21:54] there is any outliers or not there are
[21:56] different methods for this outlier uh
[21:59] removal but nowadays so those outliers
[22:04] we generally do not want to lose because
[22:06] of this fact that some of the incidences
[22:08] some of the even if we say that some of
[22:10] the natural calamities that is occurring
[22:12] that might have not we have not seen it
[22:15] before. So there are some unprecedented
[22:18] uh uh incidents that is happening. So
[22:22] that means that some of the data set may
[22:25] be there in the entire data pool which
[22:29] the the the historical data set doesn't
[22:32] support. So in in many cases of the
[22:34] statistical approach we generally uh
[22:37] tries to ignore them when we try to
[22:40] extract some of these underlying
[22:42] processes. But nowadays uh when we are
[22:46] already advanced in this computing uh e
[22:50] efforts. So both statistical and this
[22:53] machine learning generally does not
[22:55] consider throwing away any information.
[22:57] So that's definitely not a good uh good
[23:00] practice for both statistical and
[23:02] machine learning approaches and uh this
[23:05] is also not a practice in the
[23:06] statistical methods but sometimes as I
[23:09] mentioned but sometimes some of the
[23:12] simpler application when we want to when
[23:15] we just go for it is uh in case of the
[23:19] statistical method it is it it is done
[23:21] but I again once again I I repeat that
[23:24] in the general perspective if there is
[23:27] No uh issue related to the computational
[23:30] ability then we should consider all the
[23:33] data po data points. We should not uh
[23:36] lose any information.
[23:39] Selection of the predictors based on the
[23:41] correlation I this correlation again is
[23:43] a uh statistical term and sometimes we
[23:46] need to be use this term very cautiously
[23:49] rather the associ association which may
[23:52] covers both the linear or nonlinear
[23:54] things. So the based on either
[23:57] correlation or association with respect
[23:59] to the response variable that is a y
[24:02] that we discuss may lead to some better
[24:05] performance of this machine learning
[24:07] models. That means I am just
[24:10] shortlisting the uh the predictor
[24:13] variables which are quote unquote maybe
[24:15] better or more informative for my target
[24:18] uh variable here. However, it is always
[24:22] recommended that some physical or the
[24:25] conceptual justification
[24:27] uh should be there while selecting these
[24:30] predictors and this is for both the
[24:32] statistical as well as machine learning
[24:35] approaches.
[24:39] Here in a very uh small uh description
[24:42] we want to show you that the statistical
[24:45] and machine learning models in the
[24:47] perspective of their parameter and the
[24:49] sample size when we are dealing with
[24:51] this uh uh these two domains. So when we
[24:55] talk about this machine learning then
[24:58] generally the number of parameters and
[25:00] sample size is very high whereas the
[25:03] these things are generally small in case
[25:05] of the statistical approaches.
[25:08] But as we just started this particular
[25:11] uh lecture that what is their
[25:13] complimentary role to each other is that
[25:16] you see if we generally they say that
[25:18] this is the overall domain is that
[25:20] artificial intelligence within that we
[25:22] have one uh subset is called the machine
[25:24] learning and within that there are some
[25:26] kernel methods random forest boosting
[25:29] all these things we'll discuss uh in
[25:31] this lecture itself and within that
[25:33] there are some uh say neural networks is
[25:36] there very recently the rib
[25:38] uh deep learning has come. So these are
[25:40] all finds a very very effective and
[25:43] potential applications is many domain
[25:47] but this role of the statistics is
[25:49] basically almost in this all these uh
[25:52] approaches whether it is the
[25:54] pre-processing or the post-processing of
[25:55] this data and sometimes we have seen if
[25:58] we have have some sort of statistical uh
[26:01] tools and this machine learning
[26:02] approaches are combined to each other of
[26:05] course they perform uh much uh better
[26:08] way that's why These two things should
[26:10] be learned hand in hand so that uh their
[26:13] performance can yield a better
[26:16] performance as compared to their
[26:18] individual uh uh methods or uh
[26:22] applications.
[26:25] So coming to the concluding remarks of
[26:28] this uh lecture that main trade-off
[26:30] between this uh statistics and this
[26:33] machine learning is the interpretability
[26:36] versus accuracy. When we talk about this
[26:38] interpretability
[26:40] uh with relatively few parameters as
[26:43] just now we discussed and this few
[26:45] predictors the statistical models are
[26:48] much more interpretable as compared to
[26:50] the machine learning methods. So say few
[26:53] decades before even the machine learning
[26:56] models sometimes called as a blackbox
[26:58] models. So here if we just take one
[27:01] example all of you might be heard this
[27:03] term this linear regression. So this is
[27:06] this gives some useful very useful
[27:08] information that uh on each predictor uh
[27:11] how each predictor variable influences
[27:13] the response variable. On the other
[27:16] hand, if we take the one classic machine
[27:18] learning approach for example the
[27:20] artificial neural network or random
[27:22] forest, these are generally run as an
[27:26] ensemble of models initialized with
[27:29] different random numbers leading to a
[27:32] huge number of parameters.
[27:35] that are often not uh interpretable uh
[27:40] for the practical problems.
[27:42] But very recently so our data sets
[27:46] whether it is simulated or observed from
[27:48] different it is increasingly it is
[27:50] increasingly larger and much more
[27:52] complex. So now the interpretability has
[27:56] become harder and harder. Of course if
[27:59] moment the number of data sets increases
[28:02] and to achieve even with the statistical
[28:05] uh model whether whether we can really
[28:08] draw some the interpretation of that
[28:10] one. On the other hand the advantage of
[28:13] the prediction accuracy that has
[28:16] attained by these machine learning
[28:17] models those are making the increasingly
[28:20] attractive to these different uh fields
[28:23] of application.
[28:24] That means uh the statistical models are
[28:27] more interpretable whereas the machine
[28:29] learning models are less but since we
[28:32] have a very huge amount of data in the
[28:36] recent days so that interpretability
[28:38] with respect to the uh uh the
[28:40] statistical model even the challenging
[28:43] but sometimes some of the uh new
[28:45] approaches like that uh physics based uh
[28:49] machine learning and a combination of
[28:52] this uh the statistics and the machine
[28:54] learning definitely will be helpful in
[28:56] both the perspectives like their
[28:58] interpretability as well as the the
[29:02] accuracy uh of so the problem that we
[29:05] are applying to with this uh thank you
[29:08] so from this lecture and in this le next
[29:11] lecture we'll see another uh uh nice
[29:14] concept of the role of the data science
[29:17] that is also a uh very new and
[29:20] attractive domain and there are some
[29:22] interlin with this uh the statistical
[29:26] machine learning and the data science
[29:28] that's overall this module. Thank you.
