-----------------------------------------------------------------
Week 5: Learning from data: regression
제5주차: 자료에서 정보 얻어내기(회귀/경향성 분석)
- Using machine learning tools to investigate your data
수집한 자료를 조사하기 위해 기계학습 도구(분석 소프트웨어)를 활용하기
- Calculating the red-shifts of distant galaxies
먼 은하의 적색편이 계산하기
-----------------------------------------------------------------
1강: 자료 가지고 학습하기
Lesson 1: Learning from Data / 한글자막
-----------------------------------------------------------------
2강: 우주의 규모, 거리측정
Lesson 2: The Cosmological Distance Scale / 한글자막
-----------------------------------------------------------------
3강: 기계 학습의 기초
Lesson 3: What is machine learning / 한글자막
-----------------------------------------------------------------
4강: 결정 트리 분류기
Lesson 4: Decision Tree Classifier / 한글자막-----------------------------------------------------------------
Lesson 5: Estimating Redshifts using Regression / 한글자막
-----------------------------------------------------------------
5주 요약
Week 5: Module Summary / 한글자막 / 영문자막
[강의대본]
[MUSIC] We started this module with the aim of measuring redshifts for a large sample of distant galaxies. A redshift defectively gives us a distance to a galaxy and so by doing this, we can map out the universe in 3D.
0:19
Measuring spectroscopic redshift is the most accurate way of doing this but there are more galaxies for which we have imaging observations than spectroscopic ones. So we used a set of galaxies with known redshifts and train the machine learning classifier to calculate red shifts for new sets of galaxies.
0:36
This is the type of problem that machine learning is perfect for. A task where we understand how to predict new redshifts, but where it's not possible to do it with a rule-based approach.
0:47
Machine learning is a fast solution that allows us to evaluate the accuracy of the results and work out the probability of the results being correct.
0:56
It can often seem like there's lots of hidden voodoo going on in machine learning. But really all we're doing is building a model based on known data and then applying that model to unknown data. The challenge is knowing the limitations of our models and how to evaluate their accuracy effectively. We'll discuss some of these issues such as overfitting in the next module.
1:18
One of the reasons decision trees are great is that the model they produce is quite intuitive for humans to understand. In fact, the decision tree was involved in the kinds of schemes scientists have been using for generations. For example, in developing taxonomies in biology.
1:34
In this case, we were able to calculate reaches the hundreds of thousands of galaxies. effectively using them as a tool to measure distance in the universe. One of the frustrations of an observational science like astronomy is we can't set up experiments in the lab. We can't control the unknown parameters to test hypotheses or use physical equipment like rulers to measure distances.
1:57
Instead we have to cleverly construct observations to collect data that will allow us to answer our questions.
2:05
In doing so, we often have to use the astronomical objects themselves as the tools in our laboratory. And the many different ways that astronomers have developed to calculate red shifts and hence distances, is perhaps one of the most amazing examples of this. [MUSIC]
댓글 없음:
댓글 쓰기