EdX

Big Data for Agri-Food: Principles and Tools (edX)

Big Data for Agri-Food: Principles and Tools (edX)

As the big data era unfolds, developments in sensor and information technologies are evolving quickly. As a result, science and businesses are yielding enormous amounts of data. Yet, to reap the actionable business solutions data can unveil, we must learn to ask the right questions. Join Wageningen Wageningen University & Research as the data team bridges the gap between the complexity of computer science and its practical application. Decipher your unsampled big data set.

Class Deals by MOOC List - Click here and see EdX's Active Discounts, Deals, and Promo Codes.

Demystify complex big data technologies
The sheer volume of a typical data set doesn’t fit on even the largest computer. And the tools that can handle big data seem too complex to grasp. To tackle these challenges, principles – such as immutability and pure functions – will help you understand big data technology. This makes big data management accessible, regardless of the programming language.
Specifically, should you scale up or scale out, how do you process big data stacks with map-reduce, using clusters, etc. In short, learn to recognise and put into practice the scalable solution that’s right for your situation. To illustrate, we will use current tools – such as Hadoop HDFS and Apache Spark – on user-friendly, hands-on examples from the agri-food sector. However, these principles can also be applied to other sectors.

Complexity of data collection and processing
Agri-food deserves special focus when it comes to choosing robust commercial data management technologies due to its inherent variability and uncertainty. Ranked the #1 university in Animal Sciences and Agriculture, Wageningen University & Research specialises in the interdisciplinarity between its knowledge domain of healthy food and living environment on the one hand and data science, artificial intelligence (AI) and robotics on the other.
Combining data from the latest sensing technologies (e.g. weather data) with machine learning/deep learning methodologies, allows us to unlock insights we didn’t have access to before. In the areas of smart farming and precision agriculture this allows us to:

  • Better manage dairy cattle by combining animal-level data on behaviour, health and feed with milk production and composition from milking machines.
  • Reduce the amount of fertilisers (nitrogen), pesticides (chemicals) and water used on crops by monitoring individual plants with a robot or drone.
  • More accurately predict crop yields on a continental scale by combining current with historic data on soil, weather patterns and crop yields.

In short, big data will allow us to bring forth effective solutions for smarter, innovative products. The possibilities are seemingly endless!

For whom?
You are a manager or researcher with a big data set on your hands, perhaps considering investing in big data tools. You’ve done some programming before, but your skills are a bit rusty. You want to learn how to effectively and efficiently manage very large datasets. This course will enable you to see and evaluate opportunities for the application of big data technologies within your domain. Enrol now.
This course has been partially supported by the European Union Horizon 2020 Research and Innovation program (Grant #810 775, “Dragon”).

What you'll learn

  • Recognize big data characteristics (volume, velocity, variety, veracity)
  • The difference between scaling up and scaling out
  • Big data principles: immutability and pure functions
  • Processing big data with map-reduce, using clusters
  • Understand technologies: distributed file systems, Hadoop
  • How dataframes and wrapper technology (Apache Spark) make life easier
  • The big data workflow and pipeline
  • How data is organized in datalakes, using lazy evaluation
  • Develop insight how to apply this to your own case

Syllabus

Module 1: Big data definition and characteristics
In module 1, you will learn how to recognize the characteristics of a big data problem in agriculture, to see where its biggest challenge lies. Should the solution focus on size, speed, various formats or uncertainty of data? Should you scale up or scale out?

Module 2: Big data principles: what are they and why do we need them
In module 2, you'll learn the principles that are required for scaling out: immutability and pure functions, and map-reduce. What are these and why do we need them?

Module 3: Bring those principles to practice
Module 3 shows you how to bring those principles into practice. You will learn what a cluster is, and how a distributed file system in a client-server architecture works, with Hadoop. You will understand why such a system is indeed scalable.

Module 4: Big data technologies that make implementation so much easier
Module 4 goes further into the application of big data technology, the “big data stack of technologies". The main message here is that if you know what you want to do, these technologies can take the work out of your hands. For example, you will see Apache Spark, a big data technology platform, that applies map-reduce for you.

Module 5: The big data workflow and pipeline; the how and why of datalakes
Module 5 dives deeper into the data. You'll learn about datalakes and why a datalake is different from a traditional database. You'll understand what a big data workflow looks like and what a pipeline is.

Go to Class
MOOC List is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Related Courses

Big Data Analytics Using Spark (edX) EdX
University of California, San Diego,UC San DiegoX

Big Data Analytics Using Spark (edX)

Learn how to analyze large datasets using Jupyter notebooks, MapReduce and Spark as a platform. In data science, data is called “big” if it cannot fit into the memory of a single standard laptop or workstation. The analysis of big datasets requires using a cluster of tens, hundreds or thousands of computers. Effectively using such clusters requires the use of distributed files systems, such as the Hadoop Distributed File System (HDFS) and corresponding computational models, such as Hadoop, MapReduce and Spark.

Dec 5th 2023
5-12 Weeks
Big Data, Hadoop, and Spark Basics (edX) EdX
IBM

Big Data, Hadoop, and Spark Basics (edX)

This course provides foundational big data practitioner knowledge and analytical skills using popular big data tools, including Hadoop and Spark. Learn and practice your big data skills hands-on. Organizations need skilled, forward-thinking Big Data practitioners who can apply their business and technical skills to unstructured data such as tweets, posts, pictures, audio files, videos, sensor data, and satellite imagery, and more, to identify behaviors and preferences of prospects, clients, competitors, and others. ****

Self Paced
Self-Paced
Knowledge Management and Big Data in Business (edX) EdX
The Hong Kong Polytechnic University,HKPolyUx

Knowledge Management and Big Data in Business (edX)

Learn why and how knowledge management and Big Data are vital to the new business era. The business landscape is changing so rapidly that traditional management, business and computing courses do not meet the needs for the next generation of workers in the business world. Most traditional methods are of a repetitive, rule-based nature and will be gradually replaced by Artificial Intelligence.

Self Paced
Self-Paced
Big Data Computing with Spark (edX) EdX
The Hong Kong University of Science and Technology - HKUST,HKUSTx

Big Data Computing with Spark (edX)

Learn the theory and gain hands-on experience of big data systems, using Spark as the exemplary platform. Big data systems such as Hadoop and Spark emerge as enabling technologies in managing massive amounts of data across hundreds or even thousands of computing nodes. Meanwhile, cloud computing platforms have made these technologies easily accessible to individuals as well as large enterprises.

Self Paced
Self-Paced
Foundations of Data Analytics (edX) EdX
The Hong Kong University of Science and Technology - HKUST,HKUSTx

Foundations of Data Analytics (edX)

Learn the fundamental techniques for data analytics and to be prepared for learning and applying more advanced big data technologies. Foundations of Data Analytics: This course will provide fundamental techniques for data analytics, including data collection, data extraction, data integration, data cleansing, and basic machine learning techniques.

Self Paced
Self-Paced
Introducción a la Ciencia de Datos y el Big Data (edX) EdX
Tecnológico de Monterrey,TecdeMonterreyX

Introducción a la Ciencia de Datos y el Big Data (edX)

Obtén un panorama general de lo que es Data Science o Ciencia de Datos y cómo aplicarla en las organizaciones. Aprende a tomar decisiones basadas en los datos. El futuro pertenece a la ciencia de datos y a quienes la entiendan. Al igual que el petróleo y el gas impulsaron las economías de los siglos XX y XXI, los datos impulsan cada vez mas la innovación y la economía global a medida que avanzamos hacia una nueva era denominada la revolución digital.

Self Paced
Self-Paced
Big Data Solutions for Social and Economic Disparities (edX) EdX
HarvardX,Harvard University

Big Data Solutions for Social and Economic Disparities (edX)

Join Harvard University Professor Raj Chetty in this online course to understand how big data can be used to measure mobility and solve social problems. What factors increase or decrease your likelihood of economic mobility? Does the neighborhood you grew up in play a part? How different is your life from the family’s life just a few streets over?

Self Paced
Self-Paced
AI skills: Introduction to Unsupervised, Deep and Reinforcement Learning (edX) EdX
Delft University of Technology,DelftX

AI skills: Introduction to Unsupervised, Deep and Reinforcement Learning (edX)

Learn the fundamentals and principal AI concepts about clustering, dimensionality reduction, reinforcement learning and deep learning to solve real-life problems. In this course you will learn the basics of several machine learning topics to help you solve real life challenges. Unsupervised learning techniques such as clustering and dimensionality reduction are useful to make sense of large and/or high dimensional datasets that are not annotated. Deep learning is a supervised learning technique that is useful to train neural networks to solve more complicated classification and regression tasks. Finally, reinforcement learning techniques can be used to train AI agents that interact with an environment.

Self Paced
Self-Paced
Introduction to Management Information Systems (MIS): A Survival Guide (edX) EdX
Universidad Carlos III de Madrid - UC3M,UC3Mx

Introduction to Management Information Systems (MIS): A Survival Guide (edX)

Gain the skills and knowledge needed to succeed in an MIS-dominated corporate world. This MIS course will cover supporting tech infrastructures (Cloud, Databases, Big Data), the MIS development/ procurement process, and the main integrated systems, ERPs, such as SAP®, Oracle® or Microsoft Dynamics Navision®, as well as their relationship with Business Process Redesign.

Self Paced
Self-Paced
Programming for Data Science (edX) EdX
University of Adelaide,AdelaideX

Programming for Data Science (edX)

Learn how to apply fundamental programming concepts, computational thinking and data analysis techniques to solve real-world data science problems. There is a rising demand for people with the skills to work with Big Data sets and this course can start you on your journey through our Big Data MicroMasters program towards a recognised credential in this highly competitive area. Using practical activities you will learn how digital technologies work and will develop your coding skills through engaging and collaborative assignments.

Self Paced
Self-Paced