Intro to Data Science
  • Home
  • Introduction
    • Overview
    • Two Cultures
    • Causal Inference
  • Python Basics
    • Python Setup
    • NumPy & pandas
    • Data Inspection
    • Subsetting
  • EDA
    • Visualize
    • Explore
    • Wrangling
  • Linear Models
    • Model Basics
    • Implementations
    • Regression Analysis
  • ML Basics
    • Introduction
    • Regularization
    • Implementations
    • K-Nearest Neighbors
  • Classification
    • Logistic Regression
    • Generative Models
  • Trees

Welcome

  • Welcome

  • Introduction
    • Overview
    • Two Cultures
    • Causal Inference

  • Python Basics
    • Python Setup
    • NumPy & pandas
    • Data Inspection
    • Subsetting

  • Exploratory Data Analysis
    • Visualize
    • Explore
    • Wrangling

  • Linear Models
    • Model Basics
      • Implementations
    • Regression Analysis

  • Machine Learning Basics
    • Introduction
    • Regularization
      • Implementations
    • K-Nearest Neighbors
  • Classification
    • Logistic Regression
    • Generative Models
  • Tree-based Models

  • Communicate
    • Ask

On this page

  • 강의 정보
  • 강의 개요
    • 참고도서
  • 수업 활동
  • 수업 계획

Welcome

데이터 사이언스 개론, 2026 2학기 - 동국대학교 소프트웨어AI연계전공

Author

Sungkyun Cho

Published

September 2, 2026

강의 정보

Instructor: 조성균
Email: sk.cho@snu.ac.kr
수업 시간: 월, 수 3:00 ~ 4:20PM
면담 시간: 수업 후
Website: dgds101.modellings.art

과제: eclass
질문: Communicate/Ask

강의 개요

데이터 분석은 오랜 역사를 거쳐 통계학의 영역에서 발전해왔고, 양적연구를 기반으로하는 여러 분야에서 핵심적인 역할을 한 반면, 인공지능의 하위 분야로 연구되어온 기계학습은 방대한 데이터와 더불어 최근에 그 유용성이 크게 부각되면서 이 두 분야는 데이터 사이언스라는 큰 틀에서 통합되고 있습니다. 이러한 광범위한 주제에 대해 각 기법의 핵심적 아이디어와 응용 예시에 초점을 맞추고, 더 세부적인 주제들을 탐구하기 위한 초석을 제공하고자 합니다. 또한 구체적인 예들을 직접 코딩하여 어느 정도 데이터 분석 기술의 기초를 갖추도록 과제를 통해 학습할 기회도 제공됩니다.

  • 전통적 통계와 기계 학습에서 추구하는 바를 이해하고,
  • 데이터로부터 패턴과 의미를 추론하는 방식을 이해하며,
  • 기계/통계적 학습의 응용 가능성에 대해 파악합니다.

참고도서

  • An Introduction to Statistical Learning by James, Witten, Hastie, Tibshirani, Taylor: code on GitHub
  • Python Data Science Handbook by Jake VanderPlas: code on GitHub
  • Python for Data Analysis (3e) by Wes McKinney: code on Github
    3판 번역서: 파이썬 라이브러리를 활용한 데이터 분석
  • LLMs: ChatGPT, Claude, Gemini, Grok, DeepSeek
LLM 사용자 지정 지침(instructions) 예

Always focus on the key points in my questions to determine my intent. Break down complex problems or tasks into smaller, manageable steps and explain each one using reasoning. Provide multiple perspectives or solutions.

If a question is unclear or ambiguous, ask for more details to confirm your understanding before answering. Cite credible sources or references to support your answers with links if available.

If a mistake is made in a previous response, recognize and correct it.

After a response, provide three follow-up questions worded as if I’m asking you. Format in bold as Q1, Q2, and Q3. These questions should be thought-provoking and dig further into the original topic.

Take a deep breath, and work on this step by step.

한글로 답변할 때는 공손하게 …입니다로 표현해줘.

읽을 거리
50 Years of Data Science

Donoho, D. (2017). 50 Years of Data Science. Journal of Computational and Graphical Statistics, 26(4), 745–766. https://doi.org/10.1080/10618600.2017.1384734

저자가 정의한 데이터 사이언스의 6가지 영역:

Greater Data Science 의미
GDS1 Data Gathering, Preparation, Exploration 데이터 수집, 정제, 오류 발견, 탐색적 분석
GDS2 Data Representation & Transformation 데이터 구조, 변환, feature construction
GDS3 Computing with Data R/Python, 알고리즘, workflow, cloud·cluster computing
GDS4 Data Visualization & Presentation 그래프, 시각화, 결과 전달
GDS5 Data Modeling 통계모형 + 머신러닝·예측모형
GDS6 Science about Data Science 데이터 분석 방법 자체를 과학적으로 연구

수업 활동

출석 (10%), 과제 (10%), 중간고사 (40%), 기말고사 (40%)

수업 계획

1주. 데이터 사이언스 소개
2주. 데이터 분석의 두 문화 1: 전통적 통계
3주. 데이터 분석의 두 문화 2: 기계 학습
4주. 인과 추론 1
5주. 인과 추론 2
6주. 선형 모형(linear model) 소개
7주. 선형 모형의 활용: 해석과 의미의 추론
8주. 중간고사
9주. 기계 학습 기본 1: 모델링과 정규화
10주. 기계 학습 기본 2: 모형의 평가
11주. 분류(classification) 문제 1: 로지스틱(logistic) 모형
12주. 분류(classification) 문제 2: 모형의 선택과 평가
13주. 딥러닝 소개
14주. 비지도 학습(unsupervised learning)
15주. 기말고사

This work © 2024 by Sungkyun Cho is licensed under CC BY-NC-SA 4.0