官术网_书友最值得收藏!

  • Data Wrangling with Python
  • Dr. Tirthajyoti Sarkar Shubhadeep Roychowdhury
  • 297字
  • 2021-06-11 13:40:25

Python for Data Wrangling

There is always a debate on whether to perform the wrangling process using an enterprise tool or by using a programming language and associated frameworks. There are many commercial, enterprise-level tools for data formatting and pre-processing that do not involve much coding on the part of the user. These examples include the following:

  • General purpose data analysis platforms such as Microsoft Excel (with add-ins)
  • Statistical discovery package such as JMP (from SAS)
  • Modeling platforms such as RapidMiner
  • Analytics platforms from niche players focusing on data wrangling, such as Trifacta, Paxata, and Alteryx

However, programming languages such as Python provide more flexibility, control, and power compared to these off-the-shelf tools.

As the volume, velocity, and variety (the three Vs of big data) of data undergo rapid changes, it is always a good idea to develop and nurture a significant amount of in-house expertise in data wrangling using fundamental programming frameworks so that an organization is not beholden to the whims and fancies of any enterprise platform for as basic a task as data wrangling:

Figure 1.2: Google trend worldwide over the last Five years

A few of the obvious advantages of using an open source, free programming paradigm such as Python for data wrangling are the following:

  • General purpose open source paradigm putting no restriction on any of the methods you can develop for the specific problem at hand
  • Great ecosystem of fast, optimized, open source libraries, focused on data analytics
  • Growing support to connect Python to every conceivable data source type
  • Easy interface to basic statistical testing and quick visualization libraries to check data quality
  • Seamless interface of the data wrangling output with advanced machine learning models

Python is the most popular language of choice of machine learning and artificial intelligence these days.

主站蜘蛛池模板: 灵丘县| 三原县| 屯留县| 涿州市| 武穴市| 望都县| 巨野县| 合川市| 体育| 江陵县| 天水市| 灵丘县| 集安市| 垦利县| 浦江县| 衡阳市| 海口市| 浦北县| 乳山市| 龙口市| 海门市| 二连浩特市| 涪陵区| 临湘市| 大埔县| 咸阳市| 阳西县| 梁平县| 乐山市| 邹城市| 嘉定区| 土默特右旗| 新丰县| 宜章县| 滦平县| 长白| 增城市| 莱阳市| 武威市| 丽水市| 冷水江市|