10+ years of proven experience in distributed systems, parallel computing, data warehousing and analytics.
6+ years of experience in large scale hadoop compute clusters and petabyte data stores.
1 year of experience using spark data stack for map reduce, sql, streaming, machine learning, and graph applications.
Hands-on programming expertise in java, python, sql (scala a plus)
Best practices of data management - data governance, data lineage, data security, authentication, and authorization
Proficient in data life cycle - appraisal, acquisition, cleansing, metadata, ingestion, storage, transformation, and visualization
Experience in oltp, olap, and any mpp (teradata, asterdata, greenplum etc.,) data engines.
Expert at hdfs and yarn configuration and optimization (gluster, quant, tachyon, grid gain is a plus)
Well versed in flume, scribe, sqoop, oozie, pig, falcon, thrift, cascade, hive, hcatalog, zookeeper
Hadoop, Architect, Data Stack, Java, Python, SQL, Teradata, Technical Architect
Job details are sourced from the employer's original posting.
Open job postingAbout the company
We recruit for product development teams as well as IT teams; we understand the difference in skill sets and personnel needed for both back-end and front-end. We have the resources and access to a global network of talent across various technologies and industries to provide your company with the right candidates you need and just when you need them