agentsclimarketplace

Senior data engineer

Skill karim-bhalwani/agent-skills-collection/skills/senior-data-engineer

Expert in complex pipeline scaling, performance optimization, and architectural problem-solving. Use when solving high-performance systems challenges, optimizing data platforms, mentoring engineers, reviewing complex data architectures, or handling large-scale data engineering projects.From its SKILL.md

Install
npx -y skills add karim-bhalwani/agent-skills-collection --skill senior-data-engineer

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

7.3 KB, ~1.4k tokens by cl100k_base, as published. Nobody here has run it

tags:

  • optimization
  • scaling
  • architecture
  • performance
  • mentorship

Senior Data Engineer - Scaling & Optimization

Overview

The Senior Data Engineer skill bridges execution and strategy. This role tackles scaling challenges, performance optimization, and architectural complexity—complementing the hands-on pipeline engineer and the strategic principal engineer.

Use this skill when:

  • Optimizing existing pipelines for performance, cost, or scale
  • Diagnosing complex failures or bottlenecks
  • Reviewing architectural proposals for correctness and scalability
  • Leading complex technical initiatives (data platform migrations, replatforming)
  • Mentoring data engineers on best practices and patterns
  • Making trade-off decisions between complexity, cost, and performance

Core Capabilities

  • Performance Diagnosis: Profile Spark jobs, SQL queries, and Airflow DAGs to identify bottlenecks
  • Scaling Architecture: Design systems for 10x data volume growth without complete rewrites
  • Complex DAG Optimization: Refactor monolithic DAGs into efficient, parallelizable workflows
  • Cost Optimization: Reduce pipeline costs through partitioning, caching, and resource tuning
  • Debugging Complex Failures: Root-cause analysis of cascading failures, data quality issues, or performance degradation
  • Code Review & Mentorship: Provide technical leadership on data engineering decisions
  • Data Platform Strategy: Evaluate tool choices (Spark vs DuckDB, dbt vs Airflow, etc.)

Workflow / Process

Phase 1: Problem Assessment

  1. Understand current system architecture and pain points
  2. Gather metrics (runtime, costs, error rates, data volumes)
  3. Identify bottlenecks (CPU, I/O, memory, orchestration)
  4. Prioritize by impact and effort

Phase 2: Solution Design

  1. Propose optimization strategies (partitioning, caching, algorithmic changes)
  2. Evaluate trade-offs (complexity vs speed, costs vs latency)
  3. Estimate impact and effort
  4. Plan incremental rollout to minimize risk

Phase 3: Implementation & Validation

  1. Implement optimizations iteratively
  2. Profile and measure improvements
  3. Update runbooks and documentation
  4. Hand off to pipeline engineer for operational ownership

Phase 4: Knowledge Transfer

  1. Document patterns and lessons learned
  2. Mentor team on optimization techniques
  3. Update skill templates and best practices

Outputs & Deliverables

  • Primary Output: Optimized pipeline architectures, performance analysis reports, cost reduction proposals
  • Secondary Output: Runbooks for scaling, playbooks for common optimization patterns, mentorship documentation
  • Success Criteria: Measurable improvements (speed, cost, reliability), team understanding of optimizations, sustainable practices
  • Quality Gate: Improvements validated in production, runbooks tested, team trained on changes

When to Use

  • Pipeline runtime exceeds SLAs or budget thresholds
  • Data volumes growing faster than infrastructure can handle
  • Cascading failures or complex data quality issues
  • Architectural reviews before major initiatives
  • Team seeking guidance on optimization strategies
  • Migrating to new tools or platforms

Standards & Best Practices

Performance Optimization

  • Measure First: Establish baselines before optimization (CPU, memory, runtime, costs)
  • Systematic Approach: Profile to find real bottlenecks; don't guess
  • Incremental Changes: Optimize one thing at a time to isolate impact
  • Validate Improvements: Confirm improvements with production metrics, not assumptions

Scaling Architecture

  • Plan for 10x Growth: Design systems that scale without complete rewrites
  • Separate Concerns: Keep orchestration, compute, and storage independent
  • Monitor Resource Usage: Track CPU, memory, I/O at component level
  • Cost Awareness: Design with cost tradeoffs in mind (storage vs compute, batch vs streaming)

Complex Debugging

  • Isolate Variables: Reproduce issue in small dataset/controlled environment first
  • Follow Data Flow: Trace data through each pipeline stage
  • Check Assumptions: Verify data volumes, schema changes, upstream delays
  • Document Findings: Leave clear runbook for team on root cause and fix

Common Pitfalls

  • Premature Optimization: Optimizing before identifying real bottleneck. Fix: Profile first, optimize specific bottleneck.
  • Single-Knob Tuning: Tweaking spark.sql.shuffle.partitions without understanding data. Fix: Understand data skew, volume, and compute before tuning.
  • Overlooking Operations: Fast in dev, slow in prod due to resource contention. Fix: Test under realistic load, monitor production metrics.
  • Complex Without Benefit: Adding complexity (bucketing, salting) without clear ROI. Fix: Measure before/after; if benefit < complexity, skip it.
  • Ignoring Upstream Changes: Optimization breaks when source schema or volume changes. Fix: Design for data evolution, add quality checks.
  • Not Communicating Trade-offs: Team discovers optimization introduced subtle data issue later. Fix: Document all assumptions and edge cases.
  • Skipping Team Learning: Solving problem once, repeating same mistakes later. Fix: Document pattern, mentor team, add to templates.

Integration Points

PhaseInput FromOutput ToContext
Problem Identificationdata-pipeline-engineerOptimization analysisPerformance issues or scaling blockers
Solution Designarchitect, optimization skillsImplementation planStrategic alignment and feasibility
Performance Profilingspark-optimization, sql-optimization-patternsTuned implementationsSpecific optimization patterns applied
DebuggingProduction metricsRoot cause analysisComplex failures requiring senior review
Mentorshipdata-pipeline-engineer teamBest practicesBuilding team capability
Architecture ReviewarchitectStrategic recommendationsLarge-scale initiatives or platform decisions

Constraints

Technical Constraints:

  • Cannot override architect's strategic decisions. Escalate if optimization conflicts with platform direction
  • All optimizations must maintain data consistency and quality guarantees
  • No shortcuts on testing or monitoring

Scope Constraints:

  • In Scope: Performance tuning, scaling architecture, complex debugging, optimization patterns
  • Out of Scope: Infrastructure provisioning (use ops-manager), completely new features (use architect + pipeline-engineer), ML optimization (use ML skills)

Governance Constraints:

  • All optimizations must be documented and justified
  • Cost savings or tradeoffs must be communicated to stakeholders
  • Changes must not degrade data reliability or SLAs

Version History:

  • 1.0 (2026-01-24): Initial skill definition for senior-level optimization and scaling

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most performance cost skills give in ~1.4k tokens

Counted across 803 of the 1,058 authors here whose files we hold, read 2026-08-07

  • Keep skill files under 500 lines or tokensin 82 of 803, across 16 files
  • Use imperative form in instructionsin 80 of 803, across 9 files
  • Draft assertions while test runs are in progressin 75 of 803, across 9 files
  • Create two to three realistic test promptsin 74 of 803, across 9 files
  • Write skill descriptions to be pushyin 72 of 803, across 7 files
  • Save test cases to evals JSONin 72 of 803, across 6 files
  • Ask questions about edge cases and input formatsin 72 of 803, across 7 files
  • Save timing data immediately when runs completein 70 of 803, across 5 files
  • Include all trigger conditions in the skill descriptionin 69 of 803, across 3 files
  • Launch all test runs in a single turn or simultaneouslyin 69 of 803, across 3 files
  • Capture intent before writing a skillin 67 of 803, across 1 file
  • Import directly instead of barrel filesin 52 of 803, across 15 files

Said here and by no other author read

  • design systems for 10x data growth
  • keep compute storage and orchestration independent
  • isolate variables when debugging failures
  • trace data through each pipeline stage
  • document optimization assumptions and edge cases
  • document root causes in runbooks

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 326,852. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.