2018년 7월 27일 금요일

5주 천문학자 아쉬쉬 마하발과 인터뷰

[커세라 강좌 소개] 자료기반 천문학(Data-Driven Astronomy)
-----------------------------------------------------------------
Week 5: Learning from data: regression
제5주차: 자료에서 정보 얻어내기(회귀/경향성 분석)
- Using machine learning tools to investigate your data
  수집한 자료를 조사하기 위해 기계학습 도구(분석 소프트웨어)를 활용하기
- Calculating the red-shifts of distant galaxies
  먼 은하의 적색편이 계산하기
-----------------------------------------------------------------
1강: 자료 가지고 학습하기
Lesson 1: Learning from Data / 한글자막
-----------------------------------------------------------------
2강: 우주의 규모, 거리측정
Lesson 2: The Cosmological Distance Scale / 한글자막
-----------------------------------------------------------------
3강: 기계 학습의 기초
Lesson 3: What is machine learning /  한글자막
-----------------------------------------------------------------
Lesson 4: Decision Tree Classifier / 한글자막
-----------------------------------------------------------------
Lesson 5: Estimating Redshifts using Regression / 한글자막
-----------------------------------------------------------------
5주 요약
Week 5: Module Summary / 한글자막 / 영문자막
-----------------------------------------------------------------
5주 천문학자 아쉬쉬 마하발과 인터뷰
Bonus Interview with Ashish Mahabal / 한글자막 / 영문자막


My name is Ashish Mahabal, I am a senior research scientist at Caltech, specifically at the Center for Data-Driven Discovery here and I have been working on large-scale sky surveys, which means that I've been using lot of mathematical and statistical techniques. And that has gotten me interested in methodology transfer, how other fields also use this, so I have also been applying some of these techniques to Earth science and health care data and some cancer - cancer research data as well. Since I came to Caltech in 1999, I started working on Big Data because the surveys that I was involved in the Palomar-Quest Survey, before that the Digitized Palomar Observatory Survey, more recently the Catalina Real-Time Transient Survey and a little bit on the Palomar Transient Factory. So what we do is that we observe large parts of the sky again and again. So essentially, this is like taking digital movies. And that has become possible only in the last several years. Until then, it was mainly that people would go out and look at small parts of the sky, specific samples and come back and study those. But these digital movies essentially give you lots of data and moreover, you can find what is changing at different levels in the universe, in our galaxy and our solar system - and also outside of our galaxy. First of all, what one needs do is that one has to make sure that the data are good quality, that there are not too many missing data points in what you have. And we mainly work - when it comes to large data from these surveys, is work on the time series of the data sets. So the time series can be very gappy. They are heteroscedastic in the sense that the error bars can vary on the same object depending on when you are observing it. And of course, the number of objects that we have vary in brightness quite a lot. So what that means is that the time series that we deal with are quite different from what the financial services people deal with, where you have very specific times when the data are taken. And so that provides new challenges, trying to - trying to figure out what objects are doing when you are not observing them. And that's most of the time, because the amount of time we observe is really small compared to the total time where the variations in the objects are taking place. Rather than study a single object which may be doing something weird, you try to do it in a statistical nature. So for instance - and I don't work on those, but let's take an example of how stars evolve. You have a star that spends its time in, say, the main sequence for a long time and then it will- it will evolve into a giant and so on. And those time scales are so long that you don't get to see them in your own lifetime. So what you instead do is observe millions of stars - and some of them are in one phase - and some of them are in the other. Similarly, when we are looking at objects that vary in brightness, consider supernova, for instance, the supernova stage would last only for a few weeks, but before that, if you had observations and if you're lucky, you can find the star that was the progenitor of it. And then by looking at the entire time series, you can try to design specific statistical features, which then you can look for in your entire data set. So once you start understanding a little bit more about the kinds of objects that you are interested in, you design these features or filters, then, that you can use across the data set to try to find more of them. And once you have a large enough sample, then you are in business. Because then you can start applying many of the standard techniques to the data set after that. Okay. So I can answer that on two different levels. One level is getting a good data set in the first place. Most surveys are designed with specific goals in mind and what that essentially means is that you are trying to go after either some low- hanging fruit, or some specific classes. So, the data in other classes also exist in that data set, but those may not have been observed optimally in order to go after those classes. And so what would be useful is you could combine different data sets to do that. And I'm also working on what is called Domain Adaptation or Model Adaptation, where you try to combine these data sets. And then that becomes very interesting. Because when - for instance, if you want to do classification, then you may find that objects that don't vary in brightness, they hugely outnumber all other classes. And within the classes that vary, there may be some classes like the flaring M stars, which would be far more than some other class. And so what that means is that the data sets are not balanced. And if they're not balanced, most techniques don't work directly on them easily. So what you need to do is then find artificial way to balance them and make sure that the technique that you are applying makes sense, because you don't want to find correlations that don't really mean anything. Because correlation is not causation and you are always going to find some correlation. So getting a good data set - order it in a way so hat you have good balanceness and there is proper meta data that tells you enough about the data set. I think those are the biggest challenges while you are pre-processing and during the process itself; making sure that you can follow-up each step with proven answer and make it reproducible. So that's the other angle of what - the challenges. So many times, what you do is that you start with simple correlations and simple visualization. And languages like Python and R are great for that. Python is becoming the workhorse for many, many things and things like scikit-learn that they have. It's lovely to just start playing around with. R has a large number of statistical libraries which have been written by statisticians. So that's the good part. And so playing around for - with a bunch of these different ones, I would say, should be one of the first things. And I would advise people to learn both of them - Python and R - because both of them have some good things and they should have them in their repertory with them,I think. Combining diverse data sets which were not taken with the same goals in mind. There are huge data sets that are out there which are - which have not been combined in that fashion and doing something like that remains a big challenge. And I hope to see more progress happening in that area. And there are many, many new tools that are coming up that are likely to help there; for instance, in the image domain. Deep learning is getting popular everywhere and there are very good tools out there to do that. But again, the basics are of physics and mathematics and so students would want to make sure that while it's easy to use online tools and simply connect them each other - to each other and do a lot of things, going back to the basic physics and statistics is something that they should keep in mind. So one thing that has been good in astronomy is that we have been good at maintaining meta data for our data sets - so data about data. When we take images, for instance, we have been using what is called the FITS format. And the FITS images have a very good header which has all kinds of information - where were the data taken and what telescope was it, what was the size of the mirror, what was the filter and what time it was taken and was the shutter open for this long or less than that and all that. Now what we find is that because of that, we have been able to build structures or the names of the columns that we use and then be able to transfer information from one data set to another easily. And the same is not true in some of the other fields; like geoscience is still good, but when it comes to health care science, for instance, then the meta data keeping there has been at least a few years behind what we astronomers have been doing. Astronomy is fantastic because you are trying to solve the origin of the universe, you are trying to figure out where did we come from, why are we here and all that. Whereas, in healthcare, especially when you work on something like cancer - early detection research network is one area I'm working in - you are trying to see how we can continue to be here longer. And so that is, in fact, rewarding. And when one sees that the same kinds of techniques can be applied, that's fantastic. Because once you take a data set and abstract it enough, then the tools that you are using are - they don't care where the data came from, so long as you are sure and you are careful about maintaining the domain knowledge and not going to, as I said before, noise levels that are too much or don't find trivial correlations. So it's highly rewarding to be able to work on these two completely different scales from the universe level to the cell level. 



5주 요약

[커세라 강좌 소개] 자료기반 천문학(Data-Driven Astronomy)
-----------------------------------------------------------------
Week 5: Learning from data: regression
제5주차: 자료에서 정보 얻어내기(회귀/경향성 분석)
- Using machine learning tools to investigate your data
  수집한 자료를 조사하기 위해 기계학습 도구(분석 소프트웨어)를 활용하기
- Calculating the red-shifts of distant galaxies
  먼 은하의 적색편이 계산하기
-----------------------------------------------------------------
1강: 자료 가지고 학습하기
Lesson 1: Learning from Data / 한글자막
-----------------------------------------------------------------
2강: 우주의 규모, 거리측정
Lesson 2: The Cosmological Distance Scale / 한글자막
-----------------------------------------------------------------
3강: 기계 학습의 기초
Lesson 3: What is machine learning /  한글자막
-----------------------------------------------------------------
Lesson 4: Decision Tree Classifier / 한글자막
-----------------------------------------------------------------
Lesson 5: Estimating Redshifts using Regression / 한글자막
-----------------------------------------------------------------
5주 요약
Week 5: Module Summary / 한글자막 / 영문자막


[강의대본]

0:00
[MUSIC] We started this module with the aim of measuring redshifts for a large sample of distant galaxies. A redshift defectively gives us a distance to a galaxy and so by doing this, we can map out the universe in 3D. 
0:19
Measuring spectroscopic redshift is the most accurate way of doing this but there are more galaxies for which we have imaging observations than spectroscopic ones. So we used a set of galaxies with known redshifts and train the machine learning classifier to calculate red shifts for new sets of galaxies. 
0:36
This is the type of problem that machine learning is perfect for. A task where we understand how to predict new redshifts, but where it's not possible to do it with a rule-based approach. 
0:47
Machine learning is a fast solution that allows us to evaluate the accuracy of the results and work out the probability of the results being correct. 
0:56
It can often seem like there's lots of hidden voodoo going on in machine learning. But really all we're doing is building a model based on known data and then applying that model to unknown data. The challenge is knowing the limitations of our models and how to evaluate their accuracy effectively. We'll discuss some of these issues such as overfitting in the next module. 
1:18
One of the reasons decision trees are great is that the model they produce is quite intuitive for humans to understand. In fact, the decision tree was involved in the kinds of schemes scientists have been using for generations. For example, in developing taxonomies in biology. 
1:34
In this case, we were able to calculate reaches the hundreds of thousands of galaxies. effectively using them as a tool to measure distance in the universe. One of the frustrations of an observational science like astronomy is we can't set up experiments in the lab. We can't control the unknown parameters to test hypotheses or use physical equipment like rulers to measure distances. 
1:57
Instead we have to cleverly construct observations to collect data that will allow us to answer our questions. 
2:05
In doing so, we often have to use the astronomical objects themselves as the tools in our laboratory. And the many different ways that astronomers have developed to calculate red shifts and hence distances, is perhaps one of the most amazing examples of this. [MUSIC] 


2018년 7월 26일 목요일

5주/5강: 회귀 결정 트리로 적색편이 측정

[커세라 강좌 소개] 자료기반 천문학(Data-Driven Astronomy)
-----------------------------------------------------------------
Week 5: Learning from data: regression
제5주차: 자료에서 정보 얻어내기(회귀/경향성 분석)
- Using machine learning tools to investigate your data
  수집한 자료를 조사하기 위해 기계학습 도구(분석 소프트웨어)를 활용하기
- Calculating the red-shifts of distant galaxies
  먼 은하의 적색편이 계산하기
-----------------------------------------------------------------
1강: 자료 가지고 학습하기
Lesson 1: Learning from Data / 한글자막
-----------------------------------------------------------------
2강: 우주의 규모, 거리측정
Lesson 2: The Cosmological Distance Scale / 한글자막
-----------------------------------------------------------------
3강: 기계 학습의 기초
Lesson 3: What is machine learning /  한글자막
-----------------------------------------------------------------
Lesson 4: Decision Tree Classifier / 한글자막
-----------------------------------------------------------------
5강: 회귀 결정 트리로 적색편이 측정
Lesson 5: Estimating Redshifts using Regression / 한글자막 / 영문자막

* 단어에 익숙해지기
- Supervised Learning (교사학습)
- Classification(분류)와 Regression(회귀): 교사학습 알고리즘
- Decision Tree Model (결정 트리 모형): 학습기 구조의 한 형태
- Decision Tree Classifier (결정 트리 분류기)
- Decision Tree Regressor (결정 트리 회귀자)

[강의대본]



[00:06] 앞서 장난감 로봇의 예를 가지고 결정 트리 분류(classification)와 회귀(regression)가 어떻게 작동하는지 살펴봤다. 이렇게 배운 지식을 실제 자료에 적용해 볼 때가 됐다. 은하에서 방출되는 전체 빛은 전자기 스펙트럼의 모든 파장에 걸쳐 분포한다. 슬로언 디지털 스카이 서베이는 모든 천체에 대해 파장 구간 300에서 1,100 나노미터에서 5개의 광학 및 1개의 적외선 필터를 사용하여 입사량(flux)을 측정한다. 천체로부터 필터의 파장에 해당하는 빛이 얼마나 많이 방출되는지 알 수 있다. 측정된 입사량은 다시 각 필터에 붙인 이름으로 빛의 세기(등급)로 환산된다.(필터의 빛 통과 특성은 선형적이지 않다)



[00:43] 지금 보는 이 그림는 25광년 떨어진 청색의 밝은 주계열 별 베가(Vega)의 스펙트럼이다. 표면온도가 약 1만도 켈빈 가량되는 아주 뜨거운 별이다. (겉보기 색깔은) 가장 밝은 빛의 파장으로 나타났다고 할 수 있다. 이번에는 이 베가가 실제로는 아주 먼 은하에 속한 별이라고 해보자. 우주의 팽창을 말해 주듯이 베가는 적색편이를 보이고 같은 스펙트럼이지만 최고점의 위치가 더 긴 파장 쪽으로 이동하여 망원경에 장착된 여러 파장대역의 필터를 통해 다르게 측정된다.



[01:17] 필터마다 측정치를 계량화 하였는데 이를 색차라 한다. 색차는 이웃한 파장대역의 필터를 통과한 빛의 평탄화한 측정비로 이웃한 필터의 밝기 등급 차와 같다. 슬로언 필터가 다섯개 이므로 4개의 색차가 나온다. 이를테면, u-g, g-r, r-i 그리고 i-z.



[01:35] 스펙트럼이 희미한(퍼진) 빛의 은하라도 기본적으로 그 은하에 속한 모든 별의 빛의 총체다. 측광 적색편이 분류의 요점은 적색편이된 은하가 적색편이가 없을 때와 비하여 다른 관측 색차를 갖게될 것이라는 점이다. 하지만 이 생각에는 근본적인 애매함이 있는데 자료에 근거한 혹은 실증적인 방법 이외에는 문제를 풀 방법이 없다는 것이다. (적색편이는 우주팽창을 의미하는데 이는 관측에 의한 발견이며 규명된 이론은 없다.) 만일 두 은하의 관측된 색차가 다르다면 그것이 적색편이로 인한 것인지 혹은 한 은하의 스펙트럼이 다른 은하와 달리 좀 특이 했던 것인지 누가 알겠는가?



[02:10] 기계학습이 이 의문에 대한 탈출구가 되어줄 것이다. (엄청난 양의 관측자료에 기반한 기계학습이 우주의 수수께끼를 풀 수 있다.) 먼저 은하 5만개 가량의 대규모 관측자료를 가져오자. 이 자료들은 분광학적으로 적색편이가 측정되고 (분광학,spectroscopy 과 측광, photometry는 현재 가장 정밀한 우주 관측방법이다) 네가지 측광 색차가 계산되었다.



[02:24] 그런다음, 특징(이번 예의 경우 색차)를 사상하는 결정 트리 회귀자(regressor)를 훈련 시킨다. ('훈련'은 관측치로부터 추출한 특징인 색차를 입력으로 받아 이를 결정해주는 사상함수 혹은 결정트리, decision tree 를 구축하는 과정이다.) 이번 예의 경우 사상의 결과(함수에서 되돌려지는 값)은 분광학에서 얻은 적색편이다. 여기에 보는 것은 3단계 결정 트리 회귀(decision tree regression)다.

[02:39] 첫째로 자료를 분할하는 결정은 u-g 색차가 0.7475보다 작거나 같을 때다. 그 아랫단계의 분류는 g-r 이 0.372보다 작거나 같을 때다. 결정이 중지되면 그 지점의 회귀 값(regression value, value=1.1859)과 평균 제곱근 오류(mean square error, mse = 0.5109) 값을 돌려준다.




[03:00] (결정 트리)모형에서 결정을 따르기는 쉬운 반면 (모형의) 구성이 매우 복잡하여 전체를 통찰하기는 상당히 어렵다. 예를들어 우리의 실험에서 가장 우수한 성능을 내는 결정 트리는 최대 깊이가 일곱차례의 결정 과정을 거치는 경우다. 그런데 우리 모형이 얼마나 정밀할까? 우리가 사용한 표본에서 모든 은하에 대해 예상(expected)적색편이 대비 실(true)적색편이의 비교를 도표로 나타내 보았다. (분광학 정보를 알고있는 은하를 훈련에 사용하였다. 훈련이 끝난 결정 트리 모형을 검증하기 위해 훈련에 사용한 은하들을 결정트리에 넣어보고 정확도를 확인해본다.) 그림에서 보는 대로 대부분 예상치가 실제 값과 일치했다. 평균자승 오차(Root Mean Square) 분포로 이들 예측의 정확도를 평가 할 수 있다. 물론 소수의 예측이 벗어난 경우도 보인다. 때로 이를 파탄적 예외(catastrophic outliers)라고 부르긴 하는데 결정트리 모형을 심각하게 외곡 시킬 수도 있기 때문이다. 이들 오류의 원인은 잘못된 분광 분류, 색차계산의 오차, 혹은 학습에 사용된 은하가 명확하지 않았던 탓일 수도 있다. 이에 대해 좀더 심도있게 따져 볼 수 있다. 하지만 현재로서 이번 최초 결과에 만족하기로 한다.



[03:54] 이제 (훈련된)결정트리 회귀자를 분광학 자료가 없는 은하를 상대로 실행시켜보자. 실행하기 전에 먼저 훈련해 놓은 모형이 얼마나 잘 작동할지 그 회귀 정확도를 염두에 둬야한다. 각 은하가 괜찮은 적색편이 값을 알려줄 수 있는 경우에 해당하는지 아니면 비극적 예외중 하나가 될지 모르기 때문이다. 이는 분명히 엄청난 문제이긴 하다. 이를 개선할 방법은 다음주 강의에서 살펴보자.

[04:21] 최고의 예측가능한 분류와 예측을 위해 기계학습 알고리즘과 그 문제에 대한 우리의 천문학 전문지식을 결합한다. 일예로, 상당수의 측광 적색편이 분류기들은 애매함을 줄이기 위해 은하 종류별 여러가지 스펙트럼 틀(spectra template)을 가지고 있다.

[04:38] 자, 이제 한발짝 물러서서 우리가 해놓은 것이 어떻게 작동하는지 보자. 단지 네개의 색차 특징과 약간 오차가 있는 미리 분리해서 보관중인 자료를(이번 연구를 위해 특별히 세심하게 마련하지 않은 일반자료를 의미함)사용 했지만 의미있는 정밀도의 적색편이 예측(결정 트리에서 알려준 적색편이 값)을 할 수 있었다. 이에 덧붙여 분류 과정이 개인 컴퓨터에서 단 몇초 만에 빠르게 끝마칠 수 있었다. 인간이 했더라면 수년이 걸렸을 일이다. 기계학습은 대량의 자료 분석에 있어서 극도로 강력한 도구이다.

5주/4강: 결정 트리 분류기

[커세라 강좌 소개] 자료기반 천문학(Data-Driven Astronomy)
-----------------------------------------------------------------
Week 5: Learning from data: regression
제5주차: 자료에서 정보 얻어내기(회귀/경향성 분석)
- Using machine learning tools to investigate your data
  수집한 자료를 조사하기 위해 기계학습 도구(분석 소프트웨어)를 활용하기
- Calculating the red-shifts of distant galaxies
  먼 은하의 적색편이 계산하기
-----------------------------------------------------------------
1강: 자료 가지고 학습하기
Lesson 1: Learning from Data / 한글자막
-----------------------------------------------------------------
2강: 우주의 규모, 거리측정
Lesson 2: The Cosmological Distance Scale / 한글자막
-----------------------------------------------------------------
3강: 기계 학습의 기초
Lesson 3: What is machine learning /  한글자막
-----------------------------------------------------------------
Lesson 4: Decision Tree Classifier / 한글자막 / 영문자막

------------------------------------------------------
{프롤로그} 분류(Classification)와 회귀(Regression)

기계 학습(Machine Learning)에서 지도학습(Supervised Learning) 부분 인용해 보자.

---[인용시작]---
지도 학습(Supervised Learning): 사람이 교사로서 각각의 입력(x)에 대해 레이블(y)을 달아놓은 데이터를 컴퓨터에 주면 컴퓨터가 그것을 학습하는 것이다. 사람이 직접 개입하므로 정확도가 높은 데이터를 사용할 수 있다는 장점이 있다. 대신에 사람이 직접 레이블을 달아야 하므로 인건비 문제가 있고, 따라서 구할 수 있는 데이터양도 적다는 문제가 있다.

- 분류(Classification): 레이블 y가 이산적(Discrete)인 경우 즉, y가 가질 수 있는 값이 [0,1,2 ..]와 같이 유한한 경우 분류, 혹은 인식 문제라고 부른다. 일상에서 가장 접하기 쉬우며, 연구가 많이 되어있고, 기업들이 가장 관심을 가지는 문제 중 하나다. 이런 문제들을 해결하기 위한 대표적인 기법들로는 로지스틱 회귀법 [5], KNN, 서포트 벡터 머신 (SVM), 의사 결정 트리 등이 있다.

- 회귀(Regression): 레이블 y가 실수인 경우 회귀문제라고 부른다. 데이터들을 쭉 뿌려놓고 이것을 가장 잘 설명하는 직선 하나 혹은 이차함수 곡선 하나를 그리고 싶을 때 회귀기능을 사용한다. 잘 생각해보면 데이터는 입력(x)와 실수 레이블(y)의 짝으로 이루어져있고, 새로운 임의의 입력(x)에 대해 y를 맞추는 것이 바로 직선 혹은 곡선이므로 기계학습 문제가 맞다. 통계학의 회귀분석 기법 중 선형회귀 기법이 이에 해당하는 대표적인 예이다.
---[인용끝]---

[강의대본]



[00:06] (이전의 강의에서) 높은 수준의 지도학습 분류(Supervised Classification) 과정에 대하여 살펴봤다. 이제 구체적인 예를 들어보기로 한다. 결정 트리(Decision Tree)는 아마 가장 이해하기 쉬운 기계학습 알고리즘 이다. 결정트리의 표현이 인간이 결정을 내릴 때 하는 논리적적인 생각과 비슷하기 때문이다. 이번 예를 위해 기계학습 교과서에도 나올만 한 고전적인 문제를 꺼내봤다.



[오늘 테니스를 칠까요?]

[00:30] 로봇 선수 '로비'가 이 질문에 결정을 내려야 하는 기로에 있다. 그 결정은 과거의 경험에서 얻은 자료로 훈련되었다. 말하자면 오늘 테니스를 칠것인가 말것인가의 결정은 다음의 네가지 요인에 달렸다.



[00:41] 맑음, 구름 또는 비로 나타내는 통상 날씨 예측, 더움, 온화함, 추움의 온도, 보통 혹은 높음의 습도, 그리고 강함 혹은 약함의 바람. 이 훈련 자료를 가지고 테니스를 쳤던 지난날의 결정을 추적해 보자.



[01:00] 예를 들어 아홉번째 날(D9)은 추웠지만 습도는 보통, 약한바람과 함께 해가 났었기에 테니스 치기 좋은 날씨였다. 반면 열네번째 날(D14)은 강한 바람과 함게 비가와서 테니스를 치지 못했다. 이런 소규모의 자료 묶음을 가지고 테니스를 칠 것인지 결정할 결정트리를 손으로 만들어 보기로 하자.



[01:19] 처음 만든 결정트리는 이렇다. 일반 날씨 전망, 그러니까 맑음, 흐림 또는 비, 이 세개의 선택지가 있다. 이 학습 자료에 근거 한다면 구름이 드리운 날에는 항상 테니스를 쳤다. 따라서 흐림에서 가지를 뻗으면 바로 최종 결론에 이른다. 바로 테니스를 치기로 분류된다.

[01:36] 만일 맑은 날이라면 결정은 습도(Humidity)에 따라 달라진다. 맑은 날 보통의 습도라면 테니스 치기 좋다. 하지만 맑은 날이라도 습도가 높으면 테니스 치기 좋지 않다.

[01:46] 끝으로 비오는날은 바람의 조건에 달렸다. 비오는 날 바람이 약하면 테니스를 쳐도 좋지만, 비오고 바람도 강하면 좋지 않다.

[01:58] 이제 앞서 본 간단한 자료를 근거로 결정트리를 직접 만들어 보자.



[02:02] 그런데 기계학습 알고리즘은 트리의 각 단계에서 어떤 속성을 주어야 하는지 알게되는 걸까? (두번째 가지치는 단계에서 'Humidity' 대신 'Temperature'를 고를 수도 있다) 트리에서 각 단계 결정을 위해 정해진 것은 없지만 학습자는 새로운 정보를 가장 우선시 하거나 오류를 최소화 하는 방향으로 선택한다. 알고리즘 마다 이런 정보의 가치(information gain)를 취급하는 저마다 기준을 가지고 있다.

[02:19] 이견의 여지가 있지만 가장 흔한 측정법을 한가지를 들어보면, 엔트로피(Entropy)는 예측 가능성(predictablity) 혹은 불확실성(uncertainty)의 정도(분포)를 계량한다. 예를 들어 만일 구름이 드리울 것이라고 예상되면 항상 테니스를 치기로 했다. 따라서 우리는 어떤 확신을 가지고 있는 셈이므로 엔트로피는 0이다(불확실 성이 매우 낮다). 만일 비올것 같은 날씨라면 테니스를 칠 확률이 60%라도 엔트로피는 1에 가깝다(불확실 성이 매우 높다). 정보의 이득(가치)은 해당 질문에 대한 답변에 따라 엔트로피가 감소하는 방향으로 측정한다. 각 속성에 대한 정보이득은 이와 같다.



[02:52] 정보이득(Information Gain)과 관련되어 먼저 어떤 의문이 드는가? 일단 알고리즘으로 트리를 구성하고 나면 이를 미지의 자료에 적용하게된다. 따라서 만일 우리 로봇 '로비'가 어느날 일어나 보니 해가 떠있고 중간정도 습도로 덥고 강한 바람이 불었다. 우리가 만든 결정모형으로는 테니스를 쳐야한다고 예측했다. 이는 바른 결정인가? 강한 바람이 부는 더운 날씨에 정말 테니스를 치고 싶은가? 제아무리 로비라 하더라도 테니스를 칠것인지 말것인지 결정을 (결정트리에) 맞겨놓지 않았던가?



[03:19] 우리가 가진 모형은 항상 학습에 사용했던 자료에 제한된다. 이 예에서 사용했던 학습자료가 너무 적었다. 하지만 이번 강좌에서 봤듯이 훈련용 자료의 부족이 현실의 기계학습의 문제다. 예측이 탄탄하려면 보통 수천개의 훈련용 사례를 활용 한다. 이번 예에서 자료들은 모두 저마다 나름의 특징을 가지고 있었다. 온도 만 해도 덥다, 온화하다, 춥다로 세가지 특징이 있다(모두 이산적, discrete 이다). 하지만 천문학에서 관측 자료는 모두 실수 값을 갖는다.(분류보다 회귀 알고리즘을 사용하는 것이 타당하지 않을까?) 이제 온도를 차갑다, 온화하다, 뜨겁다 대신 특징을 실수 값으로 매겨 보자. 온도의 범위를 섭씨 15 에서 30도로 잡고 화씨로 치면 60 에서 90 도의 범위다.



[03:57] (온도를 단순히 이산적이지 않은 연속으로 놓았더라도) 학습자는 대략 이전과 같은 과정을 따르게된다. 단, 온도를 실수값 범위에 놓고 결정을 지어야 하는 경우만 빼고 말이다. 어떤 온도와 같거나 보다 작은 조건에서 트리를 분기 할 지 감안 해야한다. 다시말해 최적 분기(optimal split) 그 결정의 정보이득(information gain)을 최대화하는 것이어야 한다.

[04:17] 지난번 강의에서 언급한 교사 학습(supervised learning)은 분류(classification)와 회귀(regression)를 모두 포함할 수 있다. 우리의 테니스 트리는 분류의 예다. 결과가 분명히 구분되는 두 가지중 하나다. 로비는 테니스를 치거나 말거나 둘중 하나를 택한다. 우리는 또한 실수값 결과 결정트리에 대해 배웠다. 그것이 우리가 다음에 해야할 일이다. 우리가 은하의 적색편이를 계산하기 위해 결정 트리 회귀를 사용한다면 지난세기 인간 컴퓨터를 놀라게 해왔던 방식이 될 것이다.

2018년 7월 24일 화요일

Science Bulletins: Sloan Digital Sky Survey—Mapping the Universe

슬로언 디지털 스카이 서베이

Sloan Digital Sky Survey
https://www.sdss.org/surveys/

https://en.wikipedia.org/wiki/Sloan_Digital_Sky_Survey
https://en.wikipedia.org/wiki/Galaxy_Zoo



[한글 자막예정]