aws | Abstract Content Factory

Spark on AWS EMR – The Missing Manual

Posted July 1, 2015 by Dan Osipov & filed under Big Data.

Apache Spark recently received top level support on Amazon Elastic MapReduce (EMR) cloud offering, joining applications such as Hadoop, Hive, Pig, HBase, Presto, and Impala. This is exciting for me, because most of my workloads run on EMR, and utilizing Spark required either standing up manual EC2 clusters, or using EMR bootstrap, which was very… Read more »

Apache Spark on EC2

Posted September 9, 2014 by Dan Osipov & filed under Big Data.

Its easy to get started with Apache Spark. You can get a template for a Scala job using the Typesafe Activator and have it running on a local cluster with a small dataset. You can also use a handy script spark_ec2 to launch an EC2 cluster as detailed in Running Spark on EC2 document. You could… Read more »

M	T	W	T	F	S	S
		1	2	3	4	5
6	7	8	9	10	11	12
13	14	15	16	17	18	19
20	21	22	23	24	25	26
27	28	29	30	31

Posts Tagged: aws

Spark on AWS EMR – The Missing Manual

Apache Spark on EC2

Consulting

Recent Posts

Posts Tagged: aws

Consulting

Recent Posts

Tags