- Data Wrangling with Python
- Dr. Tirthajyoti Sarkar Shubhadeep Roychowdhury
- 297字
- 2021-06-11 13:40:25
Python for Data Wrangling
There is always a debate on whether to perform the wrangling process using an enterprise tool or by using a programming language and associated frameworks. There are many commercial, enterprise-level tools for data formatting and pre-processing that do not involve much coding on the part of the user. These examples include the following:
- General purpose data analysis platforms such as Microsoft Excel (with add-ins)
- Statistical discovery package such as JMP (from SAS)
- Modeling platforms such as RapidMiner
- Analytics platforms from niche players focusing on data wrangling, such as Trifacta, Paxata, and Alteryx
However, programming languages such as Python provide more flexibility, control, and power compared to these off-the-shelf tools.
As the volume, velocity, and variety (the three Vs of big data) of data undergo rapid changes, it is always a good idea to develop and nurture a significant amount of in-house expertise in data wrangling using fundamental programming frameworks so that an organization is not beholden to the whims and fancies of any enterprise platform for as basic a task as data wrangling:

Figure 1.2: Google trend worldwide over the last Five years
A few of the obvious advantages of using an open source, free programming paradigm such as Python for data wrangling are the following:
- General purpose open source paradigm putting no restriction on any of the methods you can develop for the specific problem at hand
- Great ecosystem of fast, optimized, open source libraries, focused on data analytics
- Growing support to connect Python to every conceivable data source type
- Easy interface to basic statistical testing and quick visualization libraries to check data quality
- Seamless interface of the data wrangling output with advanced machine learning models
Python is the most popular language of choice of machine learning and artificial intelligence these days.
- Design for the Future
- 自動控制原理
- 教父母學會上網
- 機艙監測與主機遙控
- 基于單片機的嵌入式工程開發詳解
- Google SketchUp for Game Design:Beginner's Guide
- 內模控制及其應用
- 悟透AutoCAD 2009案例自學手冊
- 激光選區熔化3D打印技術
- 統計挖掘與機器學習:大數據預測建模和分析技術(原書第3版)
- Ansible 2 Cloud Automation Cookbook
- Java組件設計
- Unreal Development Kit Game Design Cookbook
- Hands-On Deep Learning with Go
- Deep Learning Essentials