Algoworks logo

Lead Data Engineer

Algoworks
Remote
Noida, IndiaRemote· about 1 hour ago

Straight from Algoworks’s careers page. Apply on the company site — no recruiter, no middleman.

Explore more remote Data Engineer jobs — salaries, top companies, and the latest openings.See all →

Lead Data Engineer

Location: Noida, India

Department: PE-Cloud

Experience: 10

Skills: Databricks, PySpark, Azure Data Factory

Role: Lead Data Engineer
Location: India, Remote
Experience: 10+ Years

Algoworks

About the company
Algoworks is an award-winning artificial intelligence, engineering services and experience transformation firm with offices across the United States, Europe, South America and India. We bring together a global team of engineers, architects, designers, researchers and operators united by rigor, accountability and a commitment to delivering measurable results.
For over 20 years, Algoworks has partnered with Fortune 500 organizations across the Americas, Europe and Asia to define, build and run technology that drives meaningful business outcomes. Our work combines human-centered design, engineering excellence and AI-powered capabilities to solve complex challenges with clarity and precision. Innovation, particularly in the responsible application of AI, is embedded in how teams approach problem-solving and continuous improvement.
At Algoworks, growth is continuous and closely tied to impact. Teams collaborate across geographies and disciplines, strengthening outcomes through shared insight and collective expertise. The culture values transparency, open dialogue and an environment where every voice is heard and contribution is recognized.
Through collaboration, accountability and a focus on results, Algoworks operates at the intersection of technology and people, building not only advanced systems but strong global teams that elevate performance and create lasting impact.

Follow the video below to know about us! Clipchamp

Role overview
We are seeking a hands-on Senior Data Engineer with strong expertise in Azure Databricks and Azure Data Factory to build, optimize, and maintain scalable enterprise data pipelines.
The role will focus on high-performance ETL/ELT development, Delta Lake optimization, data processing across Bronze, Silver, and curated layers, and close collaboration with DWH and reporting teams for downstream consumption.
The ideal candidate will act as a senior technical contributor, ensuring reliability, performance, data quality, and maintainability across the data platform.

Key responsibilities:

1.Pipeline Development
  • Build and maintain scalable data pipelines using Azure Databricks and Azure Data Factory.
  • Implement ingestion and transformation logic across Bronze and Silver data layers.
  • Develop batch and incremental data-processing patterns.
  • Design reliable and reusable pipeline components for enterprise workloads.
  • Monitor and troubleshoot pipeline execution and data-processing issues.

2.Curated Layer & Delta Lake Development
  • Implement hydration, merge, and upsert logic using Delta Lake.
  • Build and maintain curated datasets aligned with data quality and business requirements.
  • Handle late-arriving data and incremental updates.
  • Implement reliable data transformation and reconciliation processes.
  • Ensure curated datasets are optimized for downstream consumption.

3.Performance & Storage Optimization
  • Optimize Delta Lake tables for performance and cost efficiency.
  • Select and tune appropriate storage formats such as Parquet and Delta.
  • Apply partitioning, compaction, and file-sizing strategies.
  • Tune Spark jobs for large-scale distributed data processing.
  • Identify and resolve performance bottlenecks across data pipelines and storage layers.

4.Downstream & DWH Collaboration
  • Work closely with DWH and reporting teams to support downstream data consumption.
  • Provide optimized datasets for reporting and analytical workloads.
  • Support data validation and reconciliation with Gold-layer outputs.
  • Collaborate with downstream teams to understand data requirements and optimize delivery.
  • Ensure consistency and reliability of data consumed by reporting and analytics platforms.

5.Engineering Best Practices
  • Implement basic CI/CD practices for data pipelines.
  • Follow coding standards, documentation, and version-control practices.
  • Maintain reusable, scalable, and maintainable pipeline code.
  • Support production troubleshooting and performance tuning.
  • Participate in Agile delivery processes and technical discussions.

6.Data Quality & Production Support
  • Implement data validation and quality checks across ingestion and transformation processes.
  • Investigate data discrepancies and pipeline failures.
  • Perform root-cause analysis and implement corrective actions.
  • Support production deployments and resolve data-processing issues.
  • Maintain reliability and consistency across enterprise data pipelines.

Required technical skills and competencies:
  • Strong hands-on experience in data engineering and enterprise data platforms.
  • Strong experience building data pipelines on Azure.
  • Advanced proficiency in PySpark.
  • Hands-on experience with Azure Databricks.
  • Strong experience with Azure Data Factory.
  • Deep knowledge of Delta Lake tuning and optimization.
  • Strong understanding of storage optimization using Parquet and Delta.
  • Strong SQL skills for data transformation, validation, and reconciliation.
  • Experience working with large datasets and distributed processing.
  • Experience implementing batch and incremental processing patterns.
  • Experience with hydration, merge, and upsert logic.
  • Experience with Git and basic CI/CD pipelines.
  • Familiarity with data quality and validation techniques.
  • Experience working in Agile delivery environments.

Must have skills:
  • Azure Databricks.
  • Azure Data Factory.
  • PySpark.
  • Delta Lake.
  • SQL.
  • Data Engineering.
  • ETL/ELT.
  • Data Pipeline Development.
  • Bronze/Silver/Curated Data Layers.
  • Delta Lake Performance Optimization.
  • Spark Performance Tuning.
  • Parquet.
  • Batch and Incremental Processing.
  • Git and Version Control.
  • Strong analytical and problem-solving skills.

Good to have skills:
  • Microsoft Fabric.
  • Streaming or near real-time data pipelines.
  • Data governance tools.
  • Metadata management tools.
  • Advanced CI/CD practices.
  • Experience with Gold-layer development.
  • Experience supporting enterprise reporting and DWH platforms.

Desired attributes:
  • Strong analytical and problem-solving capabilities.
  • Ability to independently design and develop complex data pipelines.
  • Strong focus on performance, scalability, reliability, and data quality.
  • Ability to troubleshoot complex production data issues.
  • Strong attention to detail and coding discipline.
  • Good communication and collaboration skills.
  • Ability to work effectively with DWH, reporting, and cross-functional teams.
  • Proactive approach to performance optimization and continuous improvement.
  • Strong ownership of data engineering deliverables.

Interview Process
2-3 rounds of discussion.

Similar remote jobs

All Data Engineer jobs →
Algoworks logo

Algoworks

Senior Microsoft Fabric Performance & Scalability Specialist

Remote
Noida, India
✓ From careers page· about 1 hour ago
Peraton logo

Peraton

Senior Data Platform Engineer

Remote
$112k–$179k
✓ From careers page· about 2 hours ago
Peraton logo

Peraton

Senior AWS Cloud Architect (Remote)

Remote
$112k–$179k
✓ From careers page· about 2 hours ago
Peraton logo

Peraton

Senior AWS Cloud Engineer

Remote
$104k–$166k
✓ From careers page· about 2 hours ago

Discover More than 100,000 Hidden Remote Jobs Before Everyone Else

Unlock All Remote Jobs Today

Simple pricing. Big savings on Quarterly and Yearly.

Monthly Access

$19/month
  • Instant access to fresh remote jobs from 500+ companies
  • New opportunities added hourly, often 3-7 days before anywhere else
  • Advanced filtering by role type, stack, pay, and location
  • Priority customer support
Start 7-day trial — $2.95
Most Popular

Yearly Access

$59/year
  • Everything in Monthly
  • Save $169 (~74%) vs paying monthly
  • Average job search takes ~6 months - get covered for the whole journey
  • Less than the cost of one lunch per month for competitive advantage
  • Equivalent to just ~$4.92/month
Start 7-day trial — $2.95