SQL Server Big Data Clusters
The Data Virtualization, Data Lake, and AI Platform
- 261pagine
- 10 ore di lettura
Get a head-start on SQL Server 2019's impactful feature—Big Data Clusters—which integrates large volumes of non-relational data for analysis alongside relational data. This book offers an early look at Big Data Clusters based on SQL Server 2019 Release Candidate 1, helping you stay ahead in mastering this crucial capability. The feature encompasses data virtualization, distributed computing, and relational databases, creating a comprehensive AI platform across the cluster environment. You will learn to deploy, manage, and utilize Big Data Clusters, including how to merge data from the HDFS file system with data from SQL Server instances. With clear examples and use cases, this guide equips you to start working with Big Data Clusters effectively. It covers the architectural foundations involving Kubernetes, Spark, HDFS, and SQL Server on Linux, and demonstrates how to configure and deploy clusters in both on-premises and cloud environments. The book also teaches you to write Transact-SQL queries, leveraging your existing skills to analyze data from diverse sources, including Apache Spark. With its theoretical insights and practical scripts, you’ll be prepared to unlock SQL Server 2019's potential by integrating various data types into a single view for business intelligence and machine learning. You will learn to install, manage, and troubleshoot Big Data Clusters, analyze large data volumes, manage HDFS data as relational,

