Gcp expert
Skill jpantsjoha/ai-native-developer-experience/.agents/skills/gcp-expert
GCP expert guardrails — IAM least-privilege, data boundaries, cost controls, residency, and official-source validation. Trigger when designing or reviewing any Google Cloud workload, especially agents, LLMs (Vertex AI/Gemini), or multi-tenant systems.From its SKILL.md
npx -y skills add jpantsjoha/ai-native-developer-experience --skill gcp-expertAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 11 stars11 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
5.2 KB, ~1.1k tokens by cl100k_base, as published. Nobody here has run it
GCP Expert
Gemini Enterprise is the governed execution foundation — identity, tenancy, and compliance are the design constraints, not context window size.
This skill enforces the discipline that makes GCP workloads production-safe: identity, data boundaries, cost controls, and regional residency. It is not a GCP feature tour — it is a checklist of the things that cause incidents and compliance failures when skipped.
When to use
- Designing any GCP infrastructure (new or modified)
- Before deploying agents or LLM workloads to GCP
- When reviewing a Terraform plan or architecture diagram for a GCP workload
- When a system touches multiple tenants, regulated data, or crosses regional boundaries
Procedure
-
Identity and IAM — verify least-privilege for every service account and human role:
- Service accounts have only the roles required for their specific function. No
roles/editororroles/owneron service accounts. - Workload Identity Federation preferred over service account keys for GKE / Cloud Run workloads.
- Conditional Access and IAM conditions applied where granular, time-bound, or context-aware control is required.
- Service accounts have only the roles required for their specific function. No
-
Governance and resource hierarchy — confirm policy enforcement is mechanical:
- Organisation Policies constrain allowed regions, resource types, and service enablement — not just documented conventions.
- Resource hierarchy (organisation → folder → project) reflects environment and data-classification separation.
- VPC Service Controls perimeter applied to sensitive APIs (Vertex AI, BigQuery, Cloud Storage with regulated data).
- Audit logging enabled on IAM changes and data-plane operations before any data lands.
-
Data boundaries — for every data store in the design:
- What data classification does it hold (public / internal / confidential / regulated)?
- Is encryption at rest enabled with a customer-managed key (CMEK) where required?
- Are tenant boundaries enforced at the data layer, not just the application layer?
- Does data cross a project or organisation boundary? If yes, is there an explicit data-sharing agreement?
-
Data residency — for each resource:
- Is the region constrained to the required geography (e.g.
europe-west2for UK data)? - Are Organisation Policies in place to prevent accidental multi-region or global resource creation?
- For LLM / Vertex AI calls: is the endpoint regional, not global, where residency matters?
- Is the region constrained to the required geography (e.g.
-
Cost controls — for every LLM, compute, or storage resource:
- Is there a budget alert configured (at 50%, 75%, 90%, 100%)?
- Are autoscaling upper bounds set? Unbounded autoscaling is unbounded spend.
- Are Vertex AI / LLM call volumes capped or rate-limited?
- Are committed-use discounts or Spot/Preemptible instances evaluated where appropriate?
-
Network and egress — confirm:
- Private Service Connect or VPC peering used where public endpoints are avoidable.
- Egress costs estimated for cross-region or internet-bound traffic.
- Firewall rules follow default-deny with explicit allow rules.
-
Observability — confirm:
- Cloud Monitoring dashboards exist for the workload.
- Alerting policies fire on error rate, latency, and cost thresholds.
- Log sinks route to a central logging project for retention and audit.
-
Run the Adversarial Gate — common GCP failure modes: overly-permissive service accounts, no VPC-SC on Vertex AI, uncapped autoscaling, global endpoints used for residency-sensitive data, missing budget alerts, missing Org Policy on allowed locations.
Official sources — validate before you assert
- Documentation:
cloud.google.com/docs· Architecture Center (incl. the Well-Architected Framework):cloud.google.com/architecture - Live docs via MCP (verify currency before pinning): Google Cloud offers managed MCP servers for a growing service list (BigQuery, Spanner, and more) — fetch current docs instead of relying on memory.
- GitHub, foundations:
github.com/GoogleCloudPlatform/cloud-foundation-fabric·github.com/terraform-google-modules - GitHub, agent examples:
github.com/google/adk-python·github.com/google/adk-samples·github.com/GoogleCloudPlatform/generative-ai - Rule: every service/API claim cites an official doc. Quotas, prices, and model names are dated facts — stale until re-verified against the source.
Outputs
- GCP guardrail checklist (pass/fail per item)
- IAM role matrix: principal | role | scope | justification
- Data classification and boundary map
- Budget alert confirmation
- Open findings for human review
Guardrails
- No
roles/editororroles/owneron service accounts. Ever. - Residency is a constraint, not a preference. If the data has a residency requirement, enforce it with Organisation Policy, not convention.
- Budget alerts are not optional. An unmonitored LLM workload will produce a surprise invoice.
- Audit logs are evidence. Enable them before go-live, not after an incident.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most containers cloud skills give in ~1.1k tokens
Counted across 607 of the 705 authors here whose files we hold, read 2026-09-06
- Run as non-root userin 34 of 607, across 27 files
- Use multi-stage buildsin 29 of 607
- Set resource requests and limitsin 24 of 607, across 20 files
- Configure liveness and readiness probesin 18 of 607, across 14 files
- Use named volumes for persistent datain 14 of 607, across 9 files
- Pin base image versionsin 14 of 607
- Set up environment variablesin 14 of 607, across 10 files
- Pin provider versionsin 14 of 607
- Apply least privilege RBAC permissionsin 10 of 607, across 7 files
- Create a dockerignore filein 10 of 607
- Use remote state with lockingin 9 of 607
- Pin base images by digestin 9 of 607, across 8 files
Said here and by no other author read
- Verify least-privilege for every identity
- Validate API claims against official documentation
- Prefer Workload Identity Federation
- Apply VPC Service Controls perimeters
- Enable audit logging before data lands
- Enforce region constraints with Organisation Policy
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.