Senior data engineer
Skill karim-bhalwani/agent-skills-collection/skills/senior-data-engineer
Expert in complex pipeline scaling, performance optimization, and architectural problem-solving. Use when solving high-performance systems challenges, optimizing data platforms, mentoring engineers, reviewing complex data architectures, or handling large-scale data engineering projects.From its SKILL.md
npx -y skills add karim-bhalwani/agent-skills-collection --skill senior-data-engineerAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
7.3 KB, ~1.4k tokens by cl100k_base, as published. Nobody here has run it
tags:
- optimization
- scaling
- architecture
- performance
- mentorship
Senior Data Engineer - Scaling & Optimization
Overview
The Senior Data Engineer skill bridges execution and strategy. This role tackles scaling challenges, performance optimization, and architectural complexity—complementing the hands-on pipeline engineer and the strategic principal engineer.
Use this skill when:
- Optimizing existing pipelines for performance, cost, or scale
- Diagnosing complex failures or bottlenecks
- Reviewing architectural proposals for correctness and scalability
- Leading complex technical initiatives (data platform migrations, replatforming)
- Mentoring data engineers on best practices and patterns
- Making trade-off decisions between complexity, cost, and performance
Core Capabilities
- Performance Diagnosis: Profile Spark jobs, SQL queries, and Airflow DAGs to identify bottlenecks
- Scaling Architecture: Design systems for 10x data volume growth without complete rewrites
- Complex DAG Optimization: Refactor monolithic DAGs into efficient, parallelizable workflows
- Cost Optimization: Reduce pipeline costs through partitioning, caching, and resource tuning
- Debugging Complex Failures: Root-cause analysis of cascading failures, data quality issues, or performance degradation
- Code Review & Mentorship: Provide technical leadership on data engineering decisions
- Data Platform Strategy: Evaluate tool choices (Spark vs DuckDB, dbt vs Airflow, etc.)
Workflow / Process
Phase 1: Problem Assessment
- Understand current system architecture and pain points
- Gather metrics (runtime, costs, error rates, data volumes)
- Identify bottlenecks (CPU, I/O, memory, orchestration)
- Prioritize by impact and effort
Phase 2: Solution Design
- Propose optimization strategies (partitioning, caching, algorithmic changes)
- Evaluate trade-offs (complexity vs speed, costs vs latency)
- Estimate impact and effort
- Plan incremental rollout to minimize risk
Phase 3: Implementation & Validation
- Implement optimizations iteratively
- Profile and measure improvements
- Update runbooks and documentation
- Hand off to pipeline engineer for operational ownership
Phase 4: Knowledge Transfer
- Document patterns and lessons learned
- Mentor team on optimization techniques
- Update skill templates and best practices
Outputs & Deliverables
- Primary Output: Optimized pipeline architectures, performance analysis reports, cost reduction proposals
- Secondary Output: Runbooks for scaling, playbooks for common optimization patterns, mentorship documentation
- Success Criteria: Measurable improvements (speed, cost, reliability), team understanding of optimizations, sustainable practices
- Quality Gate: Improvements validated in production, runbooks tested, team trained on changes
When to Use
- Pipeline runtime exceeds SLAs or budget thresholds
- Data volumes growing faster than infrastructure can handle
- Cascading failures or complex data quality issues
- Architectural reviews before major initiatives
- Team seeking guidance on optimization strategies
- Migrating to new tools or platforms
Standards & Best Practices
Performance Optimization
- Measure First: Establish baselines before optimization (CPU, memory, runtime, costs)
- Systematic Approach: Profile to find real bottlenecks; don't guess
- Incremental Changes: Optimize one thing at a time to isolate impact
- Validate Improvements: Confirm improvements with production metrics, not assumptions
Scaling Architecture
- Plan for 10x Growth: Design systems that scale without complete rewrites
- Separate Concerns: Keep orchestration, compute, and storage independent
- Monitor Resource Usage: Track CPU, memory, I/O at component level
- Cost Awareness: Design with cost tradeoffs in mind (storage vs compute, batch vs streaming)
Complex Debugging
- Isolate Variables: Reproduce issue in small dataset/controlled environment first
- Follow Data Flow: Trace data through each pipeline stage
- Check Assumptions: Verify data volumes, schema changes, upstream delays
- Document Findings: Leave clear runbook for team on root cause and fix
Common Pitfalls
- Premature Optimization: Optimizing before identifying real bottleneck. Fix: Profile first, optimize specific bottleneck.
- Single-Knob Tuning: Tweaking
spark.sql.shuffle.partitionswithout understanding data. Fix: Understand data skew, volume, and compute before tuning. - Overlooking Operations: Fast in dev, slow in prod due to resource contention. Fix: Test under realistic load, monitor production metrics.
- Complex Without Benefit: Adding complexity (bucketing, salting) without clear ROI. Fix: Measure before/after; if benefit < complexity, skip it.
- Ignoring Upstream Changes: Optimization breaks when source schema or volume changes. Fix: Design for data evolution, add quality checks.
- Not Communicating Trade-offs: Team discovers optimization introduced subtle data issue later. Fix: Document all assumptions and edge cases.
- Skipping Team Learning: Solving problem once, repeating same mistakes later. Fix: Document pattern, mentor team, add to templates.
Integration Points
| Phase | Input From | Output To | Context |
|---|---|---|---|
| Problem Identification | data-pipeline-engineer | Optimization analysis | Performance issues or scaling blockers |
| Solution Design | architect, optimization skills | Implementation plan | Strategic alignment and feasibility |
| Performance Profiling | spark-optimization, sql-optimization-patterns | Tuned implementations | Specific optimization patterns applied |
| Debugging | Production metrics | Root cause analysis | Complex failures requiring senior review |
| Mentorship | data-pipeline-engineer team | Best practices | Building team capability |
| Architecture Review | architect | Strategic recommendations | Large-scale initiatives or platform decisions |
Constraints
Technical Constraints:
- Cannot override
architect's strategic decisions. Escalate if optimization conflicts with platform direction - All optimizations must maintain data consistency and quality guarantees
- No shortcuts on testing or monitoring
Scope Constraints:
- In Scope: Performance tuning, scaling architecture, complex debugging, optimization patterns
- Out of Scope: Infrastructure provisioning (use ops-manager), completely new features (use architect + pipeline-engineer), ML optimization (use ML skills)
Governance Constraints:
- All optimizations must be documented and justified
- Cost savings or tradeoffs must be communicated to stakeholders
- Changes must not degrade data reliability or SLAs
Version History:
- 1.0 (2026-01-24): Initial skill definition for senior-level optimization and scaling
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most performance cost skills give in ~1.4k tokens
Counted across 803 of the 1,058 authors here whose files we hold, read 2026-08-07
- Keep skill files under 500 lines or tokensin 82 of 803, across 16 files
- Use imperative form in instructionsin 80 of 803, across 9 files
- Draft assertions while test runs are in progressin 75 of 803, across 9 files
- Create two to three realistic test promptsin 74 of 803, across 9 files
- Write skill descriptions to be pushyin 72 of 803, across 7 files
- Save test cases to evals JSONin 72 of 803, across 6 files
- Ask questions about edge cases and input formatsin 72 of 803, across 7 files
- Save timing data immediately when runs completein 70 of 803, across 5 files
- Include all trigger conditions in the skill descriptionin 69 of 803, across 3 files
- Launch all test runs in a single turn or simultaneouslyin 69 of 803, across 3 files
- Capture intent before writing a skillin 67 of 803, across 1 file
- Import directly instead of barrel filesin 52 of 803, across 15 files
Said here and by no other author read
- design systems for 10x data growth
- keep compute storage and orchestration independent
- isolate variables when debugging failures
- trace data through each pipeline stage
- document optimization assumptions and edge cases
- document root causes in runbooks
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.