Create a free account to unlock full content!
By registering, you agree to our Privacy Statement and Terms and Conditions.
| Program start date | Application deadline |
| 2026-09-01 | - |
| 2027-09-01 | - |
Program Overview
COSC 4397 Natural Language Processing
Overview
This is an introductory natural language processing course (NLP) intended for developing foundations in NLP and text mining. The broader goal is to understand how NLP tasks are carried out in the real world and how to build tools for solving practical NLP and text mining problems. Throughout the course, emphasis will be placed on understanding NLP concepts and tying NLP techniques to specific real-world applications through hands-on experience. The course covers fundamental topics in statistics and important topics in NLP such as embeddings, semantics, part of speech tagging, parsing, information retrieval, sentiment analysis, and psycholinguistics.
Administrative Details
- Syllabus
- Instructor office hours: Monday 2:30-3:30, Wednesday 2:30-5
- Teaching Assistant (TA) office hours:
- Navid Ayoobi: Monday 10-12
- Carl Aguinaldo: Thursday 8:30-9:30 PM, Friday 10 AM-12 PM
Prerequisites
The course requires a basic background in mathematics and sufficient programming skills. It will be helpful if you have taken and done well in one or more of the equivalent courses/topics such as Data Structures, Algorithms, Artificial Intelligence, Numerical methods, Data Science, or have some background in probability/statistics. The course reviews and covers required mathematical and statistical foundations. Sufficient experience for building projects in a high-level programming language (e.g., C++, Python, Java) is required.
Reading Materials
- Textbooks:
- Natural Language Processing in Python, NLTK
- SI: Statistical Inference, Casella and Berger
- FSNLP: Foundations of Statistical Natural Language Processing, Chris Manning and Hinrich Schütze
- WDM: Web Data Mining: Exploring Hyperlinks, Contents, and Usage Data, Bing Liu
- SLP: Speech and Language Processing, Jurafsky and Martin
- IR: Introduction to Information Retrieval, Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze
- Required Reference Materials:
- Online resources per topic as appearing in the schedule
- Lecture notes
Course Materials
- Slides
- Lecture notes
- Pretrained Sentence BERT embeddings for Project
- Project Rubrics, Walkthrough, Data in pickle and Excel formats
- NLP tools (POS Tagger, Chunker, Naive Bayes, etc. in Java) and templates with linked libraries for research project
Assignment Due Dates and Grading
| Component | Contribution | Due Date |
|---|---|---|
| HW1 | 5% | 9/9 |
| HW2 | 10% | 9/29 |
| HW3 | 7% | 10/15 |
| HW4 | 8% | 10/27 |
| Project | 40% | 11/30 |
| Exam 1 | 30% | 12/3 |
Rules and Policies
- Late Assignments: Late assignments will not, in general, be accepted. They will never be accepted if the student has not made special arrangements with the instructor at least one day before the assignment is due.
- Cheating: All submitted work (code, homeworks, exams, etc.) must be your own. If evidence of code sharing is found, you will receive an F grade in the course.
- Statute of Limitations: Grading questions or complaints will, in general, not be attended to beyond one week after the item in question has been returned.
Schedule of Topics
| Topic(s) | Resources |
|---|---|
| Introduction | Lecture notes/slides, Chapter 1 FSNLP, Boolean retrieval slides by H. Schütze |
| Statistical Foundations I: Basics | Lecture notes/slides, Chapter 2 FSNLP, Chapter 1 SI |
| Statistical Foundations II: Random Variables and Distributions | Lecture notes/slides, Chapter 2 SI, Chapter 3 SI |
| Words | Chapter 5 FSNLP, Chapter 6 FSNLP, OR07, OR08, OR09 |
| Markov Models and POS Tagging | Chapter 9 FSNLP, Chapter 10 FSNLP, OR10 |
| Grammar and Parsing | Lecture notes/slides, Chapter 9, 10, and 12 of SLP |
| Text Clustering and Topic Models | Lecture notes/slides/Programming resources, Derivation and Java implementation by G. Heinrich |
| Text Categorization | Chapter 3 WDM, F. Keller's tutorial on Naive Bayes |
| Advanced Topics: Neural Text Models | Lecture notes/slides, Word2Vec demo using Gensim, Word2Vec demo using Keras |
| Sentiment Analysis and Psycholinguistics | Lecture notes + slides + selected topics from Chapter 11, WDM, Paper on opinion spam: [Ott et al., 2011] |
