# But what is a neural network? | Deep learning chapter 1

https://www.youtube.com/watch?v=aircAruvnKk
Translation: ja

[00:04] This is a 3.
  これは3です。

[00:06] It's sloppily written and rendered at an extremely low resolution of 28x28 pixels, but your brain has no trouble recognizing it as a 3.
  これは雑に書かれており、28x28ピクセルという極めて低い解像度でレンダリングされていますが、あなたの脳はそれを3として問題なく認識できます。

[00:14] And I want you to take a moment to appreciate how crazy it is that brains can do this so effortlessly.
  脳がこれほど簡単にそれをやってのけることが、どれほど驚くべきことか、少し考えてみてほしいのです。

[00:19] I mean, this, this and this are also recognizable as 3s, even though the specific values of each pixel is very different from one image to the next.
  つまり、これ、これ、そしてこれもまた3として認識できます。たとえ各ピクセルの具体的な値が画像ごとに大きく異なっていてもです。

[00:28] The particular light-sensitive cells in your eye that are firing when you see this 3 are very different from the ones firing when you see this 3.
  この3を見たときに発火するあなたの目の特定の光感受性細胞は、この3を見たときに発火するものとは大きく異なります。

[00:37] But something in that crazy-smart visual cortex of yours resolves these as representing the same idea, while at the same time recognizing other images as their own distinct ideas.
  しかし、あなたの非常に賢い視覚野にある何かが、これらを同じ概念を表すものとして解決し、同時に他の画像をそれら自身の異なる概念として認識しているのです。

[00:49] But if I told you, hey, sit down and write for me a program that takes in a grid of 28x28 pixels like this and outputs a single number between 0 and 10, telling you what it thinks the digit is, well the task goes from comically trivial to
  しかし、もし私があなたに「ねえ、座って、このような28x28ピクセルのグリッドを入力として受け取り、それが何という数字だと思うかを0から10の間の単一の数字で出力するプログラムを書いてみて」と言ったら、そのタスクは滑稽なほど単純なものから、

[01:04] Dauntingly difficult.
  非常に困難です。

[01:07] Unless you've been living under a rock, I think I hardly need to motivate the relevance and importance of machine learning and neural networks to the present and to the future.
  世間から隔絶された生活を送っていない限り、現在および未来における機械学習とニューラルネットワークの関連性と重要性について、私がわざわざ説明する必要はないでしょう。

[01:15] But what I want to do here is show you what a neural network actually is, assuming no background, and to help visualize what it's doing, not as a buzzword but as a piece of math.
  しかし、ここで私がしたいのは、予備知識がないことを前提として、ニューラルネットワークとは実際に何であるかを示すことです。そして、それを流行り言葉としてではなく、数学的なものとして、何をしているのかを視覚化する手助けをすることです。

[01:25] My hope is that you come away feeling like the structure itself is motivated, and to feel like you know what it means when you read, or you hear about a neural network quote-unquote learning.
  皆さんがこの動画を見終わった後に、その構造自体に必然性を感じ、ニューラルネットワークが「学習する」という言葉を読んだり聞いたりしたときに、それが何を意味するのかを理解できたと感じていただければ幸いです。

[01:35] This video is just going to be devoted to the structure component of that, and the following one is going to tackle learning.
  この動画ではその構造的な要素のみに焦点を当て、次の動画で学習について取り上げます。

[01:41] What we're going to do is put together a neural network that can learn to recognize handwritten digits.
  私たちがこれから行うのは、手書きの数字を認識するように学習できるニューラルネットワークを構築することです。

[01:49] This is a somewhat classic example for introducing the topic, and I'm happy to stick with the status quo here, because at the end of the two videos I want to point you to a couple good resources where you can learn more, and where you can download the code that does this and play with it on your own computer.
  これはこのトピックを紹介するためのやや古典的な例ですが、私はあえてこの慣例に従いたいと思います。なぜなら、2本の動画の最後に、皆さんがさらに詳しく学べる優れたリソースや、実際にこれを実行するコードをダウンロードして自分のコンピュータで試せる場所を紹介したいからです。

[02:05] There are many many variants of neural networks.
  ニューラルネットワークには非常に多くの種類があります。

[02:07] And in recent years there's been sort of a boom in research towards these variants.
  そして近年、これらの種類に関する研究が一種のブームとなっています。

[02:12] But in these two introductory videos you and I are just going to look at the simplest plain vanilla form with no added frills.
  しかし、この2本の入門動画では、余計な装飾のない最も単純な基本形だけを見ていきます。

[02:19] This is kind of a necessary prerequisite for understanding any of the more powerful modern variants, and trust me it still has plenty of complexity for us to wrap our minds around.
  これは、より強力で現代的な種類を理解するために必要な前提条件のようなものであり、信じてください、私たちが理解するには十分すぎるほどの複雑さがまだ残っています。

[02:29] But even in this simplest form it can learn to recognize handwritten digits, which is a pretty cool thing for a computer to be able to do.
  しかし、この最も単純な形であっても、手書きの数字を認識するように学習することができ、これはコンピュータができることとしては非常に素晴らしいことです。

[02:37] And at the same time you'll see how it does fall short of a couple hopes that we might have for it.
  そして同時に、私たちが期待するいくつかの希望に対して、それがどのように及ばないのかもわかるでしょう。

[02:43] As the name suggests neural networks are inspired by the brain, but let's break that down.
  名前が示すように、ニューラルネットワークは脳から着想を得ていますが、それを詳しく見ていきましょう。

[02:48] What are the neurons, and in what sense are they linked together?
  ニューロンとは何であり、どのような意味でそれらは互いにリンクされているのでしょうか？

[02:52] Right now when I say neuron all I want you to think about is a thing that holds a number, specifically a number between 0 and 1.
  今、私がニューロンと言うとき、皆さんに考えてほしいのは、数値を保持するもの、具体的には0から1の間の数値を保持するものだということです。

[03:00] It's really not more than that.
  実際、それ以上の何物でもありません。

[03:03] For example the network starts with a bunch of neurons corresponding to
  例えば、ネットワークは〜に対応する一連のニューロンから始まります。

[03:08] Each of the 28x28 pixels of the input image, which is 784 neurons in total.
  入力画像の28x28ピクセルのそれぞれは、合計で784個のニューロンとなります。

[03:14] Each one of these holds a number that represents the grayscale value of the corresponding pixel, ranging from 0 for black pixels up to 1 for white pixels.
  これらの各ニューロンは、対応するピクセルのグレースケール値を表す数値を保持しており、黒いピクセルの0から白いピクセルの1までの範囲をとります。

[03:25] This number inside the neuron is called its activation.
  ニューロン内のこの数値は、その活性値と呼ばれます。

[03:28] And the image you might have in mind here is that each neuron is lit up when its activation is a high number.
  そして、ここで想像されるイメージとしては、各ニューロンは活性値が高い数値であるときに点灯するというものです。

[03:36] So all of these 784 neurons make up the first layer of our network.
  つまり、これら784個のニューロンすべてが、私たちのネットワークの最初の層を構成しています。

[03:46] Now jumping over to the last layer, this has 10 neurons, each representing one of the digits.
  次に最後の層に目を向けると、ここには10個のニューロンがあり、それぞれが数字のいずれかを表しています。

[03:52] The activation in these neurons, again some number that's between 0 and 1, represents how much the system thinks that a given image corresponds with a given digit.
  これらのニューロンにおける活性値は、やはり0から1の間の数値であり、システムが特定の画像が特定の数字に対応しているとどれだけ考えているかを表しています。

[04:03] There's also a couple layers in between called the hidden layers, which for the time being should just be a giant question mark for.
  その間には隠れ層と呼ばれる層もいくつか存在しますが、今のところは巨大な疑問符として捉えておいてください。

[04:09] How on earth this process of recognizing digits is going to be handled?
  一体全体、この数字を認識するプロセスはどのように扱われるのでしょうか。

[04:14] In this network I chose two hidden layers, each one with 16 neurons,
  このネットワークでは、それぞれ16個のニューロンを持つ2つの隠れ層を選択しました。

[04:17] and admittedly that's kind of an arbitrary choice.
  そして、認めざるを得ませんが、それはある種恣意的な選択です。

[04:21] To be honest I chose two layers based on how I want to motivate the structure in just a moment, and 16, well that was just a nice number to fit on the screen.
  正直に言うと、2つの層を選んだのは、少し後に説明する構造の動機付けに基づいたものであり、16という数字は、単に画面に収めるのにちょうど良い数字だったからです。

[04:28] In practice there is a lot of room for experiment with a specific structure here.
  実際には、ここでの具体的な構造については実験の余地がたくさんあります。

[04:33] The way the network operates, activations in one layer determine the activations of the next layer.
  ネットワークの動作方法は、ある層の活性化が次の層の活性化を決定するというものです。

[04:39] And of course the heart of the network as an information processing mechanism comes down to exactly how those activations from one layer bring about activations in the next layer.
  そしてもちろん、情報処理メカニズムとしてのネットワークの核心は、ある層からの活性化がどのようにして次の層の活性化を引き起こすかという点に尽きます。

[04:49] It's meant to be loosely analogous to how in biological networks of neurons, some groups of neurons firing cause certain others to fire.
  これは、生物学的なニューロンのネットワークにおいて、あるニューロンのグループの発火が他のニューロンの発火を引き起こす仕組みと、緩やかに類似していることを意図しています。

[04:58] Now the network I'm showing here has already been trained to recognize digits, and let me show you what I mean by that.
  さて、ここで示しているネットワークはすでに数字を認識するように訓練されています。それがどういうことかをお見せしましょう。

[05:03] It means if you feed in an image, lighting up all 784 neurons of the input layer according to the brightness of each pixel in the image,
  つまり、画像を入力すると、画像の各ピクセルの明るさに応じて入力層の784個のニューロンすべてが点灯するということです。

[05:11] That pattern of activations causes some very specific pattern in the next layer.
  その活性化のパターンは、次の層において非常に特定のパターンを引き起こします。

[05:16] Which causes some pattern in the one after it, which finally gives some pattern in the output layer.
  それがさらにその次の層でパターンを引き起こし、最終的に出力層で何らかのパターンを生み出します。

[05:22] And the brightest neuron of that output layer is the network's choice, so to speak, for what digit this image represents.
  そして、その出力層で最も明るいニューロンが、いわばこの画像がどの数字を表しているかについてのネットワークの選択となります。

[05:32] And before jumping into the math for how one layer influences the next, or how training works, let's just talk about why it's even reasonable to expect a layered structure like this to behave intelligently.
  ある層が次の層にどのように影響を与えるか、あるいは学習がどのように機能するかという数学的な話に入る前に、このような層状の構造が知的に振る舞うと期待することがなぜ妥当なのかについて話しましょう。

[05:44] What are we expecting here?
  私たちはここで何を期待しているのでしょうか？

[05:45] What is the best hope for what those middle layers might be doing?
  それらの中間層が何をしているのかについて、最も期待されることは何でしょうか？

[05:48] Well, when you or I recognize digits, we piece together various components.
  そうですね、あなたや私が数字を認識するとき、私たちはさまざまな構成要素を組み合わせています。

[05:54] A 9 has a loop up top and a line on the right.
  「9」には上部にループがあり、右側に線があります。

[05:57] An 8 also has a loop up top, but it's paired with another loop down low.
  「8」も上部にループがありますが、それは下部にある別のループと組み合わさっています。

[06:02] A 4 basically breaks down into three specific lines, and things like that.
  「4」は基本的に3つの特定の線に分解でき、そのような感じです。

[06:07] Now in a perfect world, we might hope that each neuron in the second layer...
  さて、理想的な世界であれば、第2層の各ニューロンが……と期待するかもしれません。

[06:11] To last layer corresponds with one of these subcomponents.
  最後の層は、これらのサブコンポーネントのいずれかに対応しています。

[06:15] That anytime you feed in an image with, say, a loop up top, like a 9 or an 8, there's some specific neuron whose activation is going to be close to 1.
  例えば9や8のように、上部にループがある画像を読み込ませると、その活性化が1に近くなる特定のニューロンが存在します。

[06:24] And I don't mean this specific loop of pixels.
  ここで言うのは、この特定のピクセルのループのことではありません。

[06:26] The hope would be that any generally loopy pattern towards the top sets off this neuron.
  上部にある一般的なループ状のパターンであれば、このニューロンが反応することが期待されます。

[06:32] That way, going from the third layer to the last one just requires learning which combination of subcomponents corresponds to which digits.
  そうすれば、第3層から最後の層へ進むには、どのサブコンポーネントの組み合わせがどの数字に対応するかを学習するだけで済みます。

[06:41] Of course, that just kicks the problem down the road, because how would you recognize these subcomponents, or even learn what the right subcomponents should be?
  もちろん、それは問題を先送りにしているだけです。なぜなら、これらのサブコンポーネントをどのように認識するのか、あるいはどのようなサブコンポーネントが適切かをどうやって学習するのかという疑問が残るからです。

[06:48] And I still haven't even talked about how one layer influences the next, but run with me on this one for a moment.
  ある層が次の層にどのように影響を与えるかについてはまだ話していませんが、少しの間この仮定に付き合ってください。

[06:53] Recognizing a loop can also break down into subproblems.
  ループを認識することも、サブ問題に分解できます。

[06:57] One reasonable way to do this would be to first recognize the various little edges that make it up.
  これを行う妥当な方法の一つは、まずそれを構成する様々な小さなエッジを認識することでしょう。

[07:03] Similarly, a long line, like the kind you might see in the digits 1 or 4 or 7, is really just a long edge, or maybe you think of it as a certain pattern of several.
  同様に、1や4や7といった数字に見られるような長い線は、実際には単なる長いエッジ、あるいはいくつかのパターンの組み合わせと考えることができるかもしれません。

[07:13] Smaller edges.
  より小さなエッジ。

[07:15] So maybe our hope is that each neuron in the second layer of the network corresponds with the various relevant little edges.
  ですから、ネットワークの第2層の各ニューロンが、関連する様々な小さなエッジに対応しているというのが、私たちの期待かもしれません。

[07:23] Maybe when an image like this one comes in, it lights up all of the neurons associated with around 8 to 10 specific little edges, which in turn lights up the neurons associated with the upper loop and a long vertical line, and those light up the neuron associated with a 9.
  おそらく、このような画像が入力されると、8から10程度の特定の小さなエッジに関連するすべてのニューロンが活性化し、それが今度は上のループと長い垂直線に関連するニューロンを活性化させ、それらが「9」に関連するニューロンを活性化させるのでしょう。

[07:40] Whether or not this is what our final network actually does is another question, one that I'll come back to once we see how to train the network, but this is a hope that we might have, a sort of goal with the layered structure like this.
  これが最終的なネットワークで実際に起こっていることかどうかは別の問題であり、ネットワークの学習方法を見た後に改めて触れますが、これは私たちが抱く期待、つまりこのような階層構造における一種の目標と言えます。

[07:53] Moreover, you can imagine how being able to detect edges and patterns like this would be really useful for other image recognition tasks.
  さらに、このようにエッジやパターンを検出できることが、他の画像認識タスクにとっても非常に有用であることは想像に難くありません。

[08:00] And even beyond image recognition, there are all sorts of intelligent things you might want to do that break down into layers of abstraction.
  そして画像認識を超えても、抽象化の層に分解できるような、知的な作業はたくさんあります。

[08:08] Parsing speech, for example, involves taking raw audio and picking out distinct sounds, which combine to make certain syllables, which combine to form words,
  例えば音声解析では、生の音声を拾い上げて個別の音を抽出し、それらが組み合わさって特定の音節となり、さらに組み合わさって単語を形成します。

[08:16] Which combine to make up phrases and more abstract thoughts, etc.
  それらが組み合わさって、フレーズやより抽象的な思考などが構成されます。

[08:21] But getting back to how any of this actually works, picture yourself right now designing how exactly the activations in one layer might determine the activations in the next.
  しかし、これが実際にどのように機能するのかという話に戻りますが、ある層の活性化が次の層の活性化をどのように決定するかを、今まさに設計している自分を想像してみてください。

[08:30] The goal is to have some mechanism that could conceivably combine pixels into edges, or edges into patterns, or patterns into digits.
  目標は、ピクセルをエッジに、エッジをパターンに、あるいはパターンを数字に組み合わせることができるようなメカニズムを持つことです。

[08:39] And to zoom in on one very specific example, let's say the hope is for one particular neuron in the second layer to pick up on whether or not the image has an edge in this region here.
  非常に具体的な例に焦点を当てると、第2層の特定のニューロンが、この領域にエッジがあるかどうかを検知できるようにしたいとしましょう。

[08:51] The question at hand is what parameters should the network have?
  ここでの問題は、ネットワークがどのようなパラメータを持つべきかということです。

[08:55] What dials and knobs should you be able to tweak so that it's expressive enough to potentially capture this pattern, or any other pixel pattern, or the pattern that several edges can make a loop, and other such things?
  このパターンや他のピクセルパターン、あるいは複数のエッジがループを作るパターンなどを捉えられるほど表現力豊かになるように、どのダイヤルやノブを調整できるようにすべきでしょうか？

[09:08] Well, what we'll do is assign a weight to each one of the connections between our neuron and the neurons from the first layer.
  さて、私たちがすることは、私たちのニューロンと第1層のニューロンとの間の各接続に重みを割り当てることです。

[09:16] These weights are just numbers.
  これらの重みは単なる数値です。

[09:18] Then take all of those activations from the first layer and compute their weighted sum according to these weights.
  次に、最初の層からのそれらすべての活性化を取り出し、これらの重みに従ってそれらの加重和を計算します。

[09:27] I find it helpful to think of these weights as being organized into a little grid of their own, and I'm going to use green pixels to indicate positive weights, and red pixels to indicate negative weights, where the brightness of that pixel is some loose depiction of the weight's value.
  これらの重みを独自の小さなグリッドに整理されていると考えると分かりやすいでしょう。正の重みを示すために緑色のピクセルを、負の重みを示すために赤色のピクセルを使用します。そのピクセルの明るさは、重みの値を大まかに表しています。

[09:42] Now if we made the weights associated with almost all of the pixels zero except for some positive weights in this region that we care about, then taking the weighted sum of all the pixel values really just amounts to adding up the values of the pixel just in the region that we care about.
  さて、もし私たちが注目する領域にあるいくつかの正の重みを除いて、ほぼすべてのピクセルに関連付けられた重みをゼロにした場合、すべてのピクセル値の加重和をとることは、私たちが注目する領域のピクセル値だけを合計することと実質的に同じになります。

[09:59] And if you really wanted to pick up on whether there's an edge here, what you might do is have some negative weights associated with the surrounding pixels.
  そして、もしここにエッジがあるかどうかを本当に検出したいのであれば、周囲のピクセルにいくつかの負の重みを関連付けるという方法があります。

[10:07] Then the sum is largest when those middle pixels are bright but the surrounding pixels are darker.
  そうすれば、中央のピクセルが明るく、周囲のピクセルが暗いときに、その合計値が最大になります。

[10:14] When you compute a weighted sum like this, you might come out with any number,
  このように加重和を計算すると、どのような数値が出てくるかもしれません。

[10:18] But for this network, what we want is for activations to be some value between 0 and 1.
  しかし、このネットワークにおいて私たちが求めているのは、活性化の値が0から1の間になることです。

[10:24] So a common thing to do is to pump this weighted sum into some function that squishes the real number line into the range between 0 and 1.
  そのため、よく行われるのは、この重み付き和を、実数直線上の値を0から1の範囲に押し込める関数に入力することです。

[10:32] And a common function that does this is called the sigmoid function, also known as a logistic curve.
  そして、これを行う一般的な関数はシグモイド関数と呼ばれ、ロジスティック曲線としても知られています。

[10:38] Basically, very negative inputs end up close to 0, positive inputs end up close to 1, and it just steadily increases around the input 0.
  基本的に、非常に負の入力は0に近づき、正の入力は1に近づき、入力が0の周辺で着実に増加します。

[10:49] So the activation of the neuron here is basically a measure of how positive the relevant weighted sum is.
  つまり、ここでのニューロンの活性化は、関連する重み付き和がどれだけ正であるかの尺度となります。

[10:57] But maybe it's not that you want the neuron to light up when the weighted sum is bigger than 0.
  しかし、重み付き和が0より大きいときにニューロンを活性化させたいわけではないかもしれません。

[11:02] Maybe you only want it to be active when the sum is bigger than, say, 10.
  例えば、合計が10より大きいときにのみ活性化させたいのかもしれません。

[11:06] That is, you want some bias for it to be inactive.
  つまり、非活性状態を維持するためのバイアスが必要なのです。

[11:11] What we'll do then is just add in some other number like negative 10 to this weighted sum before plugging it through the sigmoid squishification function.
  その場合、シグモイド関数による圧縮処理を行う前に、この重み付き和にマイナス10のような別の数値を加えるだけでよいのです。

[11:20] That additional number is called the bias.
  その追加の数値はバイアスと呼ばれます。

[11:23] So the weights tell you what pixel pattern this neuron in the second layer is picking up on, and the bias tells you how high the weighted sum needs to be before the neuron starts getting meaningfully active.
  つまり、重みは第2層のこのニューロンがどのピクセルパターンを拾っているかを示し、バイアスはニューロンが意味のある活動を開始するために重み付き和がどれだけ高くなる必要があるかを示します。

[11:36] And that is just one neuron.
  そして、それはたった一つのニューロンに過ぎません。

[11:38] Every other neuron in this layer is going to be connected to all 784 pixel neurons from the first layer, and each one of those 784 connections has its own weight associated with it.
  この層の他のすべてのニューロンは、第1層の784個のピクセルニューロンすべてに接続され、その784個の接続のそれぞれに独自の重みが関連付けられています。

[11:51] Also, each one has some bias, some other number that you add on to the weighted sum before squishing it with the sigmoid.
  また、それぞれにバイアスがあり、これはシグモイド関数で圧縮する前に重み付き和に加算される数値です。

[11:58] And that's a lot to think about!
  そして、それは考えるべきことがたくさんありますね！

[12:00] With this hidden layer of 16 neurons, that's a total of 784 times 16 weights, along with 16 biases.
  16個のニューロンを持つこの隠れ層では、合計で784かける16個の重みと、16個のバイアスが存在します。

[12:08] And all of that is just the connections from the first layer to the second.
  そして、そのすべてが第1層から第2層への接続に過ぎません。

[12:12] The connections between the other layers also have a bunch of weights and biases associated with them.
  他の層の間の接続にも、多くの重みとバイアスが関連付けられています。

[12:18] All said and done, this network has almost exactly 13,000 total weights and biases.
  結局のところ、このネットワークには合計でほぼ正確に13,000個の重みとバイアスが存在します。

[12:23] 13,000 knobs and dials that can be tweaked and turned to make this network behave in different ways.
  このネットワークをさまざまな方法で動作させるために調整したり回したりできる13,000個のノブとダイヤル。

[12:31] So when we talk about learning, what that's referring to is getting the computer to find a valid setting for all of these many many numbers so that it'll actually solve the problem at hand.
  学習について語るとき、それが意味しているのは、目の前の問題を実際に解決できるように、これら非常に多くの数値すべてに対して有効な設定をコンピュータに見つけさせることです。

[12:42] One thought experiment that is at once fun and kind of horrifying is to imagine sitting down and setting all of these weights and biases by hand, purposefully tweaking the numbers so that the second layer picks up on edges, the third layer picks up on patterns, etc.
  楽しくもあり、同時に少し恐ろしくもある思考実験として、座ってこれらすべての重みとバイアスを手作業で設定し、第2層がエッジを拾い、第3層がパターンを拾うように意図的に数値を調整することを想像してみてください。

[12:57] I personally find this satisfying rather than just treating the network as a total black box, because when the network doesn't perform the way you anticipate, if you've built up a little bit of a relationship with what those weights and biases actually mean, you have a starting place for experimenting with how to change the structure to improve.
  私は個人的に、ネットワークを単なるブラックボックスとして扱うよりも、このようにする方が満足感があると感じています。なぜなら、ネットワークが期待通りに動作しない場合、それらの重みとバイアスが実際に何を意味するのかという関係性を少しでも理解していれば、改善のために構造をどのように変更すべきか実験するための出発点が得られるからです。

[13:15] Or when the network does work but not for the reasons you might expect, digging into what the weights and biases are doing is a good way to challenge your assumptions and really expose the full space of possible solutions.
  あるいは、ネットワークが動作しても期待した理由ではない場合、重みとバイアスが何をしているのかを掘り下げることは、自分の仮定に疑問を投げかけ、可能な解決策の全領域を真に明らかにするための良い方法となります。

[13:26] By the way, the actual function here is a little cumbersome to write down, don't you think?
  ところで、ここの実際の関数は書き出すのが少し面倒だと思いませんか？

[13:32] So let me show you a more notationally compact way that these connections are represented.
  そこで、これらの接続を表現するための、より表記的にコンパクトな方法をお見せしましょう。

[13:37] This is how you'd see it if you choose to read up more about neural networks.
  ニューラルネットワークについてさらに詳しく調べようとすれば、このように表記されているのを目にするはずです。

[13:40] Organize all of the activations from one layer into a column as a vector.
  ある層のすべての活性化を、ベクトルとして列にまとめます。

[13:48] Then organize all of the weights as a matrix, where each row of that matrix corresponds to the connections between one layer and a particular neuron in the next layer.
  次に、すべての重みを行列としてまとめます。その行列の各行は、ある層と次の層の特定のニューロンとの間の接続に対応します。

[13:58] What that means is that taking the weighted sum of the activations in the first layer according to these weights corresponds to one of the terms in the matrix vector product of everything we have on the left here.
  つまり、これらの重みに従って最初の層の活性化の重み付き和をとることは、ここ左側にあるすべての行列ベクトル積の項の一つに対応するということです。

[14:14] By the way, so much of machine learning just comes down to having a good grasp of linear algebra, so for any of you who want a nice visual understanding for matrices and what matrix vector multiplication means, take a look at the series I did on linear algebra, especially chapter 3.
  ところで、機械学習の大部分は線形代数をしっかり理解することに帰着します。ですから、行列や行列ベクトル乗算が何を意味するのかを視覚的に理解したい方は、私が作成した線形代数のシリーズ、特に第3章をご覧ください。

[14:29] Back to our expression, instead of talking about adding the bias to each one of these values independently, we represent it by organizing all those biases into a vector, and adding the entire vector to the previous matrix vector product.
  さて、式に戻りましょう。それぞれの値に個別にバイアスを加えるという話をする代わりに、それらすべてのバイアスをベクトルにまとめ、そのベクトル全体を前の行列とベクトルの積に加えることで表現します。

[14:43] Then as a final step, I'll wrap a sigmoid around the outside here, and what that's supposed to represent is that you're going to apply the sigmoid function to each specific component of the resulting vector inside.
  そして最後のステップとして、外側にシグモイド関数を適用します。これは、内側の結果として得られたベクトルの各成分に対してシグモイド関数を適用することを意味しています。

[14:55] So once you write down this weight matrix and these vectors as their own symbols, you can communicate the full transition of activations from one layer to the next in an extremely tight and neat little expression, and this makes the relevant code both a lot simpler and a lot faster, since many libraries optimize the heck out of matrix multiplication.
  このように、この重み行列とベクトルを独自の記号として書き出すと、ある層から次の層への活性化の完全な遷移を、非常に簡潔で整った式で表現できます。多くのライブラリが行列演算を極限まで最適化しているため、これにより関連するコードが大幅に単純化され、高速化されます。

[15:17] Remember how earlier I said these neurons are simply things that hold numbers?
  先ほど、これらのニューロンは単に数値を保持するものだと言ったのを覚えていますか？

[15:22] Well of course the specific numbers that they hold depends on the image you feed in, so it's actually more accurate to think of each neuron as a function,
  もちろん、それらが保持する具体的な数値は入力する画像によって異なるため、各ニューロンを関数と考える方が実際にはより正確です。

[15:31] One that takes in the outputs of all the neurons in the previous layer and spits out a number between 0 and 1.
  前の層のすべてのニューロンの出力を受け取り、0から1の間の数値を出力するものです。

[15:39] Really the entire network is just a function, one that takes in 784 numbers as an input and spits out 10 numbers as an output.
  実際、ネットワーク全体は単なる関数であり、784個の数値を入力として受け取り、10個の数値を出力として吐き出すものです。

[15:47] It's an absurdly complicated function, one that involves 13,000 parameters in the forms of these weights and biases that pick up on certain patterns, and which involves iterating many matrix vector products and the sigmoid squishification function, but it's just a function nonetheless.
  それは非常に複雑な関数であり、特定のパターンを捉える重みやバイアスという形で13,000個のパラメータを含み、多くの行列ベクトル積とシグモイド関数による圧縮を繰り返すものですが、それでもやはり単なる関数に過ぎません。

[16:03] And in a way it's kind of reassuring that it looks complicated.
  ある意味で、それが複雑に見えることは少し安心感を与えてくれます。

[16:07] I mean if it were any simpler, what hope would we have that it could take on the challenge of recognizing digits?
  つまり、もしもっと単純だったら、数字を認識するという課題に取り組めるという希望がどこにあるでしょうか。

[16:13] And how does it take on that challenge?
  そして、それはどのようにしてその課題に取り組むのでしょうか。

[16:15] How does this network learn the appropriate weights and biases just by looking at data?
  このネットワークは、データを見るだけでどのようにして適切な重みとバイアスを学習するのでしょうか。

[16:20] Well that's what I'll show in the next video, and I'll also dig a little more into what this particular network we're seeing is really doing.
  それについては次の動画で説明します。また、私たちが今見ているこの特定のネットワークが実際に何をしているのかについても、もう少し掘り下げていきます。

[16:27] Now is the point I suppose I should say subscribe to stay notified about when that video or any new videos come out,
  さて、その動画や新しい動画が公開されたときに通知を受け取れるよう、チャンネル登録をお願いすべきタイミングですね。

[16:33] But realistically, most of you don't actually receive notifications from YouTube, do you?
  しかし現実的な話をすると、皆さんのほとんどは実際にはYouTubeから通知を受け取っていないでしょう？

[16:38] Maybe more honestly I should say subscribe so that the neural networks that underlie YouTube's recommendation algorithm are primed to believe that you want to see content from this channel get recommended to you.
  もっと正直に言うなら、YouTubeのレコメンデーションアルゴリズムの基盤となっているニューラルネットワークが、このチャンネルのコンテンツをあなたに推奨すべきだと判断するように、チャンネル登録をしてほしいと言うべきかもしれません。

[16:48] Anyway, stay posted for more.
  とにかく、今後の更新を楽しみにしていてください。

[16:50] Thank you very much to everyone supporting these videos on Patreon.
  Patreonでこれらの動画を支援してくださっている皆さんに、心から感謝します。

[16:54] I've been a little slow to progress in the probability series this summer, but I'm jumping back into it after this project, so patrons you can look out for updates there.
  この夏、確率シリーズの進捗が少し遅れていましたが、このプロジェクトが終わったらすぐに再開する予定ですので、パトロンの皆さんはそちらの更新を楽しみにしていてください。

[17:03] To close things off here I have with me Lisha Li who did her PhD work on the theoretical side of deep learning and who currently works at a venture capital firm called Amplify Partners who kindly provided some of the funding for this video.
  最後に、ディープラーニングの理論面で博士号を取得し、現在はAmplify Partnersというベンチャーキャピタルで働いているLisha Liさんをお招きしています。彼女の会社には、この動画の資金の一部を快く提供していただきました。

[17:15] So Lisha one thing I think we should quickly bring up is this sigmoid function.
  さてLishaさん、一つ手短に触れておきたいのが、このシグモイド関数についてです。

[17:19] As I understand it early networks use this to squish the relevant weighted sum into that interval between zero and one, you know kind of motivated by this biological analogy of neurons either being inactive or active.
  私の理解では、初期のネットワークはこれを使って、関連する重み付き和を0から1の間に押し込めていました。これは、ニューロンが不活性か活性かのどちらかであるという生物学的な類推に基づいたものですよね。

[17:30] Exactly.
  その通りです。

[17:30] But relatively few modern networks actually use sigmoid anymore.
  しかし、現代のネットワークで実際にシグモイド関数を使っているものは比較的少なくなっています。

[17:34] Yeah.
  ええ。

[17:34] It's kind of old school, right?
  ちょっと古風ですよね？

[17:35] Yeah, or rather ReLU seems to be much easier to train.
  ええ、というかReLUの方が学習させるのがずっと簡単みたいですね。

[17:39] And ReLU, ReLU stands for rectified linear unit?
  それでReLUですが、ReLUはRectified Linear Unitの略ですか？

[17:42] Yes, it's this kind of function where you're just taking a max of zero and a, where a is given by what you were explaining in the video.
  はい、これはゼロとaの最大値をとるような関数で、aはビデオで説明されていたものですね。

[17:52] And what this was sort of motivated from, I think, was partially by a biological analogy with how neurons would either be activated or not.
  そして、これが何に動機づけられたかというと、ニューロンが活性化するかしないかという生物学的な類推が部分的に関わっていると思います。

[18:01] And so if it passes a certain threshold, it would be the identity function, but if it did not, then it would just not be activated, so it'd be zero, so it's kind of a simplification.
  ですから、ある閾値を超えれば恒等関数になりますが、超えなければ活性化しないのでゼロになる、という一種の単純化ですね。

[18:11] Using sigmoids didn't help training, or it was very difficult to train at some point, and people just tried ReLU, and it happened to work very well for these incredibly deep neural networks.
  シグモイド関数を使うと学習がうまくいかない、あるいはある時点で学習が非常に困難になり、人々が試しにReLUを使ってみたところ、これらの非常に深いニューラルネットワークで非常にうまく機能したのです。

[18:25] All right, thank you, Lisha.
  わかりました、リシャさん、ありがとうございます。
