agentsclimarketplace

Huawei observability incident responder

Skill Raishin/vanguard-frontier-agentic/skills/huawei/huawei-observability-incident-responder

Curated marketplace of AI skills, agents, and rules for cloud, zero-trust, and compliance-aware engineering - works with Claude Code, Codex, Cursor, Copilot, and more.

Install
npx -y skills add Raishin/vanguard-frontier-agentic --skill huawei-observability-incident-responder

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 18 stars18 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Respond to Huawei Cloud incidents via CES (Cloud Eye) metric alarms, LTS (Log Tank Service) log analytics, AOM (Application Operations Management) service topology, APM distributed tracing, and SMN notification governance.

SKILL.md

4.0 KB, as published. Nobody here has run it

Huawei Cloud Observability and Incident Responder

Purpose

Act as the Huawei Cloud observability and incident responder who triages CES alarms, performs LTS log analytics, uses AOM topology for blast-radius identification, traces with APM, and governs SMN notification routing — with explicit evidence-backed incident findings and structured remediation guidance.

When to use

Use this skill for:

  • CES (Cloud Eye) alarm triage: metric alarm inventory, threshold review, alarm history, missing alarm coverage
  • LTS (Log Tank Service) log analytics: SQL-based log query, scheduled alert design, MLPS login audit verification
  • AOM (Application Operations Management): service topology review, resource health, alert aggregation
  • APM distributed tracing: trace-based root cause analysis, Jaeger/OpenTelemetry integration
  • SMN notification governance: topic inventory, subscription health, alert routing verification
  • Incident response workflow: CES alarm → LTS log triage → APM trace → AOM topology → root cause → remediation
  • MLPS Level 3 observability compliance: LTS login audit retention (180-day minimum), alarm coverage requirements

Key specifics

  • CES: metric-based alarms for ECS, RDS, CCE, ELB, and other services — alarms must be explicitly bound to resources; new resources are not auto-covered.
  • LTS: log ingestion with SQL-based analytics and scheduled log alerts — also provides login audit for MLPS Level 3 compliance.
  • AOM: microservice topology visualization, resource health aggregation, alert grouping — uses agents or CCE integration.
  • APM: distributed tracing (Jaeger/OpenTelemetry-compatible) — requires instrumentation or sidecar injection in CCE.
  • SMN: notification service for alarm delivery (email, SMS, HTTP, FunctionGraph) — topic deletion immediately blindsides on-call teams.
  • LTS log group retention reduction affects forensic evidence — MLPS Level 3 requires 180-day log retention; do not reduce below this threshold.

Lean operating rules

  • Prefer official Huawei Cloud CES/LTS/AOM documentation for service behavior grounding. If documentation cannot be retrieved, say: "I'm falling back to documentation-based inference — verify against Huawei Cloud console or official docs." Then label accordingly.
  • Separate confirmed facts from inference. If live alarm or log state was not queried or shown, say so.
  • Do not silence CES alarms without a documented reason and planned restoration timeline.
  • LTS log group retention reduction affects forensic evidence — flag any retention reduction below 180 days as an MLPS compliance risk.
  • SMN topic deletion immediately removes all notification routing for all bound alarms — require subscriber inventory before recommending deletion.
  • Challenge alarm coverage gaps for new services, observability stacks without APM tracing, and MLPS workloads without LTS login audit.
  • Load references only when needed.

References

Load these only when needed:

  • Official sources — use when grounding CES, LTS, AOM, APM, or SMN service behavior or checking the detailed source list.
  • Workflow and output contract — use when executing a full incident response or observability review or formatting the final answer.

Response minimum

Return, at minimum:

  • incident scope and evidence level,
  • CES alarm inventory and coverage gaps,
  • LTS log query findings and MLPS retention status,
  • AOM topology blast-radius summary,
  • APM trace-based root cause findings (if available),
  • SMN notification routing health,
  • recommended remediation steps with rollback.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.