In recent years, the term “big data” has been hyped to the extreme, as if you were out of touch if you didn’t talk about big data.
After cloud computing, big data has become a hot trend in the IT industry. Harvard Business Review even declared that “data scientist is the sexiest job of the 21st century.” So-called “sexy” implies an irresistible allure, but also a mystique that few truly understand.
Data Science, as a broad emerging career path, is full of opportunities. The market’s demand for data talent is increasingly intense. Data scientists are in huge demand not only in the US and Europe; according to McKinsey data, there is a global talent shortage of over 200,000 for this role.
Many universities have started offering specialized programs in data analytics. Data Science has become an increasingly competitive major for applicants in recent years. But precisely because of this, herd mentality is common—many students lack clear self-awareness and decide to apply just because it’s trendy. Such hasty decisions can lead to an awkward situation.
So in this 4W1H column, I’ve once again invited my super-talented buddy, Teacher Han, to give everyone a primer on what kind of person is suitable for studying DS and what kind of development awaits after studying DS. So you can take fewer detours and make the right choices. (My network is unimaginably broad!)
![]()
- In 2013, earned dual B.S. degrees in Applied Mathematics and Physics from Stony Brook University in three years, graduating with honors (Magna Cum Laude).
- After undergraduate, entered NYU’s Courant Institute of Mathematical Sciences – the top-ranked program in the U.S. – to pursue a master’s in Data Science within the Scientific Computing track. Assisted renowned Chinese-American mathematician Professor Deng Yuefan in compiling and publishing “Lectures, Problems and Solutions for Ordinary Differential Equations” with World Scientific Publishing.
- During summer breaks in undergrad and master’s, served as a research assistant at the National Supercomputer Center in Jinan and the National Supercomputer Center in Guangzhou at Sun Yat-sen University. After completing his master’s in September 2015, has been pursuing a Ph.D. in Applied Mathematics at Stony Brook University.
Why Data Science?
Q: With a strong background in math and physics, why didn’t you continue in those directions? What sparked your interest in Data Science?
A: Math and physics are foundational disciplines that provide the theoretical basis for practical applications. They serve as foundations because many fields rely on this knowledge. Conversely, these subjects offer a rich array of future development options, which was a key reason I chose to study them.
Through my undergraduate studies, I found that my interest in applied fields outweighed my passion for pure theory. So I shifted my focus toward applied areas, using applied mathematics as an entry point to explore specific directions.
Applied math mainly has three areas: computational math, financial math, and statistics. Computational math is relatively theory-heavy; finance wasn’t my strong suit; that left statistics. It was right when DS was emerging, and academia widely believed it would be a major future direction. NYU had just launched its DS program. This interdisciplinary field combining statistics, math, and computer science caught my eye—by serendipity, I rode an academic wave.
Q: Based on your understanding, what exactly does Data Science study? How does it differ from statistics?
A: In my view, DS is a combination of statistics, math, and computer science. From statistics, DS involves probability distributions, statistical inference, linear regression; from math, it involves linear algebra; from CS, it involves programming and algorithms. Simply put, DS tells us how many ways we can process experimental data after we get it, what conclusions each method yields, and what those conclusions mean.
On the surface, DS and statistics have many commonalities—they both learn probability and statistical inference, for instance. But there are some differences.
First, statistics emphasizes theoretical foundations more than DS. For example, hypothesis testing in statistical inference is a key and difficult part in statistics, whereas in DS you just need to know its forms and calculation steps.
Second, statistics requires not only knowing how to use hypothesis testing for data analysis but also learning how to design experiments to get desired data.
Third, for statistics, mature tools exist to directly run statistical processing, while in DS you don’t necessarily need algorithms to calculate constraints. DS extends statistical methods with many algorithms to compute relationships within data.
Additionally, DS has diverse algorithms that often require coding from scratch, whereas statistics has many ready-made tools for direct processing.
Q: What skills did you cultivate and improve during your graduate study at NYU?
A: First, all my professors at NYU assumed that students had already mastered the necessary programming skills. Regardless of whether a DS course was offered by the CS department, all in-class teaching was theoretical; how to put things into practice was left for students to explore independently through homework.
I remember the first assignment required Python, and at that time I knew nothing about Python. I spent two full days self-learning, then jumped straight into using it for the homework. Second, the workload was so heavy there was no time to breathe, especially in the first two weeks while learning Python. Although each assignment gave us a week, I could barely finish on time with help from classmates and TAs. Yet on submission day, a new wave of assignments hit. This intense hands-on process constantly reinforced every concept I learned, and my coding ability improved rapidly through continuous practice.
Worth noting: all assignments had a structured system, revolving around a problem that progressed layer by layer, culminating in a global conclusion at the end. This approach helped integrate theory according to practical operation sequences. Such exercises repeated many times built up sensitivity to various data types. When encountering a problem again, I could clearly dissect the corresponding logic and algorithms from head to toe—this is a crucial ability I think.
Q: Many students interested in the data field don’t know where to start. Any advice for beginners?
A: Still considering three angles: statistics, math, and computer science. For statistics, in my view, basic statistics is built on concepts of discrete and continuous probability distribution models. So for beginners, thoroughly mastering these probability distribution models will greatly help deeper learning later. Next is linear regression.
Math: In my personal opinion, DS uses more of a “mathematical mindset” than statistics does. I emphasize mindset rather than concepts. That is, to make certain breakthroughs in this field, broadly speaking, the definitions introduced in linear algebra—like vectors, matrices, spaces—foster a strong “algebraic” thinking mode during learning, which is crucial for DS.
Finally, computer science: For this, I’d recommend learning Python directly. If you have a background in systematic CS learning, that will definitely help with picking up new languages. If you have no coding experience, Python is still a good choice: it’s open source with piles of resources, and many DS-related libraries are already mature. Also, I don’t suggest starting algorithm preparation purely from a CS perspective. From my own learning experience, maybe begin from the math or statistics angle, and after building some foundation, then incorporate algorithm concepts related to CS.
One more point: for those who already have some background in the above three areas, I think deeply understanding parallel computing concepts is very important. Because processing big data on a single machine is gradually becoming impractical—you need more powerful computing resources. But for those with major gaps, this can be postponed until later.
Q: Any application advice for students? Any recommended programs?
A: Have coding skills before enrollment, though this varies based on background. Python and R are foundational programming languages for data science. Beyond these, mastering other languages like SQL, MATLAB, and C++ can be listed on your resume as bonus skills.
Related internship experience is best; if not, do some relevant projects. Seek help from your current professors to find opportunities to join such projects. These practical experiences not only boost applications but also help you adapt more quickly to study and research in this field. Browse DS forums or groups to understand developments—this might help with your PS. Quantifying how much it helps is difficult, but after all, you’re entering this industry; knowing its frontiers and trends is excellent preparation. The application process has its own procedures and metrics. As I understand, the fundamental reason an applicant gets admitted is meeting the school’s expectations and requirements: they think you have great passion for the field and can handle the coursework. Of course, there are still quantitative thresholds at admission. The three points above not only strengthen your application but, more importantly—putting applications aside—will also build your own knowledge base.
Q: Having lived in New York for so many years, tell us about your experience.
A: For me, New York is more of a regional concept than just administrative divisions. I lived at Stony Brook—on the north shore of central Long Island—for about 5 years, in Queens for 1 year, in New Jersey for 1 year, and studied in Manhattan for 2 years. New York is roughly the collective name for my daily life areas.
When at NYU, classrooms were all in the city—literally downtown—with so many great eateries that I still don’t know where the NYU dining hall is. What I ate most was McDonald’s, the one right by the QR line. Usually in the evening, I’d grab a meal to go and get on the subway home. I couldn’t enjoy the good food because of mountains of homework; every day was a deadline, not exaggerated. The subway ride was about 80 minutes, slightly shorter after moving to New Jersey. On the subway was the most relaxing time—feeling I didn’t have to face endless classes, endless assignments, or eat the exact same fried chicken with more batter than meat. When the subway ran on time and I found a seat, it was the luckiest thing that day.
About three or four times a month, I’d bask in the sun in the garden formed by the Courant and Stern buildings. Only at those moments would I occasionally think about experiences, feelings, or reflections. That, to me, is the most genuine feeling of New York. Maybe moments of pause and reflection were extremely rare, but I could sense I was genuinely doing something meaningful—though articulating what that meaning is would be hard. But later, when I left Manhattan and returned to Stony Brook, I sometimes still missed that place. After all, that Vietnamese beef noodle shop nearby—plenty of meat, rich broth—was truly delicious.
What is Data Science?
Data science is about extracting information and knowledge from data—it’s an extension of data mining and predictive analytics, a process of discovering knowledge from data. So, in plain terms, data science analyzes data to mine potential insights hidden in that data.
Using vast amounts of data to support business decision making (data driven decision making) is the ultimate goal of data science. Broadly, research and application directions in Data Science are divided into:
Predictive Analytics:
Analyzing data to forecast what might happen in the future.
Descriptive Analytics:
Analyzing data to identify characteristics of past events and trends of ongoing events.
Prescriptive Analytics
Analyzing data to find the best course of action and achieve optimized results.
A clearer picture of the learning content can be seen from the curriculum. Judging from the course arrangements, the difficulty is not low, placing emphasis on building students’ abilities in computer science, math/statistics/data mining, and data visualization skills.
Which school to choose?

Above is a list of some Data Science master’s programs compiled by the consultant team, covering schools and programs of different tiers, including standalone DS programs and DS-related tracks within larger departments like Systems. This is for your reference.
Next, let’s take Columbia University as an example for a detailed introduction: Columbia University’s MS in Data Science Track.
Columbia has a world-class Institute for Data Sciences and Engineering where students participate in experiments and research projects. This program was first offered in Fall 2014, showing the university’s emphasis on this field. Housed in the Engineering School, the program focuses on data mining, algorithms, and statistical modeling (e.g., Algorithms for Data Science, Machine Learning for Data Science, Statistical Inference & Modeling) with a curriculum tailored to industry needs.
The following are the main courses:

Among the seven required courses, the focus is on Computer Science and Statistics. The curriculum looks rigorous and substantial; the Capstone Project also provides great opportunities for practice and application. Coupled with Columbia’s Ivy League prestige and the engineering school’s overall teaching quality and resources, the program’s competitive advantage is considerable.
Application requirements:
Like most DS programs, there are certain requirements for CS and math background—foundational computer language and math courses must be met. Currently, only Fall admission is open.
Refer to 2016 admission data:

The patterns of top schools are obvious: despite no minimum score requirements, a quick look at the average scores is sobering. So everyone should work hard to raise your standardized test scores to boost competitiveness!
Where to go after graduate?
The greatest advantage of a Data Science master’s program lies in its curriculum—it often touches on software systems, machine learning, databases, optimization, decision science, statistics, business intelligence, etc. So compared with peers from pure statistics or CS backgrounds, students with a DS master’s have a more rational and comprehensive knowledge structure. Precisely because of this, employment prospects are broader and the outlook is very promising.
The following numbers illustrate how scarce data talent is: On LinkedIn, 36,000 data scientist positions are waiting to be filled. Another site’s data shows over 6,000 companies were recruiting data talent by the end of last year. A data scientist with a Ph.D. can easily command a six-figure starting salary; after two years on the job, they can earn $200,000–$300,000 per year. It’s fair to say data talent is a highly sought-after major with high pay and high employment rates.
As things stand now, Information Technology, Insurance, and Marketing/BI are the main recruiters of data scientists. Overall, the employment situation is excellent—everybody wants them.
How to apply?
Employment prospects and application competition are directly proportional. Different programs have varying background requirements and admission criteria. Let’s take my alma mater NYU’s Data Science program as an example.
Overall, most DS programs tend to admit students with backgrounds in quantitative disciplines like math or statistics, and they also hope applicants have software programming foundations and can write programs to analyze data. The more competitive the program, the truer this is.
Math foundation:
The math trifecta: calculus, linear algebra, and probability & statistics—these are core courses for science and engineering majors. Although lacking some coursework background doesn’t necessarily mean you won’t get admitted, you’ll still be at a disadvantage. If your coursework background is insufficient, business school programs focusing more on analytics rather than data science may be more suitable, as they are friendlier to applicants from various backgrounds. Some schools have special requirements; for instance, Northwestern expects applicants to have taken Java, and NCSU has a very strict interview. This adds to the application difficulty.
Standardized test scores and application materials:
An undergraduate GPA of 3.6 (many official sites say 3.0+), TOEFL/IELTS scores of 100+/7.0+ are considered competitive baselines.
NYU’s DS program is one of the few that accepts GMAT; business school programs usually accept GMAT, but most DS programs are not in business schools. So my advice is: if you want unrestricted school choices, better take the GRE.
SOP is very important! Basically, every school’s admission committee hopes to see in your essays that you have some understanding of data science/business analytics, not just blindly applying without a clue. At the same time, as a career-oriented program, relevant work experience and internships are pluses. If you have work experience, make sure to integrate it to reflect your understanding of this field. If you don’t have work experience, it’s even more important to design your essays with appropriate content to fully demonstrate that your background and foundation can handle this major. I want to add: if you absolutely hate programming, barely passed college math, or have terribly weak interpersonal skills, you can automatically filter out data science/analytics and related programs.









