Architecture design
๐ฌ A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | ็ฒพ้ 23,000+ AI Agent ๆ่ฝๅบ๏ผ่ฆ็8ๅคง็คพไผ็งๅญฆๅญฆ็ง็ๅฎ่ฏ็ ็ฉถใCoPaper.AI 20ๅ้ๅฎๆไธ็ฏๅฏๅค็ฐ็่ง่ๅฎ่ฏ่ฎบๆ๏ผๅนถๆฏๆ็จๆทไธไผ Skillsใ-- Maintained by CoPaper.AI from Stanford REAP.
npx -y skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill architecture-designAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
What its author says it does
Copied from the file, not written here
Use only when creating new registrable ML components that require Factory or Registry patterns.
SKILL.md
9.0 KB, as published. Nobody here has run it
Architecture Design - ML Project Template
This skill defines the standard code architecture for machine learning projects based on the template structure. When modifying or extending code, follow these patterns to maintain consistency.
Overview
The project follows a modular, extensible architecture with clear separation of concerns. Each module (data, model, trainer, analysis) is independently organized using factory and registry patterns for maximum flexibility.
When to Use
Use this skill when:
- Creating a new Dataset class that needs
@register_dataset - Creating a new Model class that needs
@register_model - Creating a new module directory with
__init__.pyfactory wiring - Initializing a new ML project structure from scratch
- Adding new component types such as Augmentation, CollateFunction, or Metrics
When Not to Use
Do not use this skill when:
- Modifying existing functions or methods
- Fixing bugs in existing code
- Adding helper functions or utilities
- Refactoring without adding new registrable components
- Making simple code changes to a single file
- Modifying configuration files
- Reading or understanding existing code
Key indicator: if the task does not require a @register_* decorator or a Factory pattern, skip this skill.
Core Design Patterns
Factory Pattern
Each module uses a factory to create instances dynamically:
# Example from data_module/dataset/__init__.py
DATASET_FACTORY: Dict = {}
def DatasetFactory(data_name: str):
dataset = DATASET_FACTORY.get(data_name, None)
if dataset is None:
print(f"{data_name} dataset is not implementation, use simple dataset")
dataset = DATASET_FACTORY.get('simple')
return dataset
For detailed guidance, refer to references/factory_pattern.md.
Registry Pattern
Components register themselves via decorators:
# Example from data_module/dataset/simple_dataset.py
@register_dataset("simple")
class SimpleDataset(Dataset):
def __init__(self, data):
self.data = data
For detailed guidance, refer to references/registry_pattern.md.
Auto-Import Pattern
Modules automatically discover and import submodules:
# Example from data_module/dataset/__init__.py
models_dir = os.path.dirname(__file__)
import_modules(models_dir, "src.data_module.dataset")
For detailed guidance, refer to references/auto_import.md.
Directory Structure
project/
โโโ run/
โ โโโ pipeline/ # Main workflow scripts
โ โ โโโ training/ # Training pipelines
โ โ โโโ prepare_data/ # Data preparation pipelines
โ โ โโโ analysis/ # Analysis pipelines
โ โโโ conf/ # Hydra configuration files
โ โโโ training/ # Training configs
โ โโโ dataset/ # Dataset configs
โ โโโ model/ # Model configs
โ โโโ prepare_data/ # Data prep configs
โ โโโ analysis/ # Analysis configs
โ
โโโ src/
โ โโโ data_module/ # Data processing module
โ โ โโโ dataset/ # Dataset implementations
โ โ โโโ augmentation/ # Data augmentation
โ โ โโโ collate_fn/ # Collate functions
โ โ โโโ compute_metrics/ # Metrics computation
โ โ โโโ prepare_data/ # Data preparation logic
โ โ โโโ data_func/ # Data utility functions
โ โ โโโ utils.py # Module-specific utilities
โ โ
โ โโโ model_module/ # Model implementations
โ โ โโโ brain_decoder/ # Brain decoder models
โ โ โโโ model/ # Alternative model location
โ โ
โ โโโ trainer_module/ # Training logic
โ โโโ analysis_module/ # Analysis and evaluation
โ โโโ llm/ # LLM-related code
โ โโโ utils/ # Shared utilities
โ
โโโ data/
โ โโโ raw/ # Original, immutable data
โ โโโ processed/ # Cleaned, transformed data
โ โโโ external/ # Third-party data
โ
โโโ outputs/
โ โโโ logs/ # Training and evaluation logs
โ โโโ checkpoints/ # Model checkpoints
โ โโโ tables/ # Result tables
โ โโโ figures/ # Plots and visualizations
โ
โโโ pyproject.toml # Project configuration
โโโ uv.lock # Dependency lock file
โโโ TODO.md # Task tracking
โโโ README.md # Project documentation
โโโ .gitignore # Git ignore rules
For detailed directory structure with file descriptions, refer to references/structure.md.
Module Organization
Creating a New Dataset
When adding a new dataset:
- Create file in
src/data_module/dataset/ - Use
@register_dataset("name")decorator - Inherit from
torch.utils.data.Dataset - Implement
__init__,__len__,__getitem__
from torch.utils.data import Dataset
from typing import Dict
import torch
from src.data_module.dataset import register_dataset
@register_dataset("custom")
class CustomDataset(Dataset):
def __init__(self, data):
self.data = data
def __len__(self):
return len(self.data)
def __getitem__(self, i: int) -> Dict[str, torch.Tensor]:
return self.data[i]
Creating a New Model
CRITICAL: Models use config-driven pattern
When adding a new model:
- Create file in
src/model_module/model/or appropriate module subdirectory - Use
@register_model('ModelName')decorator __init__accepts ONLYcfgparameter - all hyperparameters come from configforward()returns dict:{"loss": loss, "labels": labels, "logits": logits}- Handle training vs inference modes using
self.training
from src.model_module.brain_decoder import register_model
@register_model('MyModel')
class MyModel(nn.Module):
def __init__(self, cfg):
super().__init__()
self.cfg = cfg
self.task = cfg.dataset.task
# ALL parameters from cfg
self.hidden_dim = cfg.model.hidden_dim
self.output_dim = cfg.dataset.target_size[cfg.dataset.task]
def forward(self, x, labels=None, **kwargs):
if self.training:
# Training logic
pass
else:
# Inference logic
pass
return {"loss": loss, "labels": labels, "logits": logits}
Adding Data Augmentation
When adding augmentation:
- Create file in
src/data_module/augmentation/ - Implement transformation function
- Register with factory if needed
Code Style Guidelines
For comprehensive style guidelines, refer to references/code_style.md.
Key principles:
- Always use type hints for function signatures
- Follow import order: standard library โ third-party โ local
- Module
__init__.pyfiles contain factory/registry logic - Model classes must be config-driven
Configuration Management
The project uses Hydra for configuration management:
- Config files in
run/conf/organize by module - Each stage (training, analysis) has its own config structure
- Use YAML files for all configuration
When Working on This Project
Before Modifying Code
- Read the relevant module's factory/registry pattern
- Check existing implementations for consistency
- Follow the established directory structure
- Use registration decorators for new components
Adding New Features
- Determine which module the feature belongs to
- Check if similar functionality exists
- Follow factory/registry pattern if creating new component types
- Add configuration files if needed
- Update documentation
Code Review Checklist
- Uses factory/registry pattern appropriately
- Follows module directory structure
- Has proper type annotations
- Imports are correctly ordered
- Registration decorator is used
- Configuration files are added if needed
Additional Resources
Reference Files
For detailed information, consult:
references/structure.md- Detailed directory structure with file descriptionsreferences/factory_pattern.md- Factory pattern in-depth explanationreferences/registry_pattern.md- Registry pattern in-depth explanationreferences/auto_import.md- Auto-import pattern in-depth explanationreferences/code_style.md- Comprehensive code style guidelines
Example Files
Working examples in examples/:
examples/custom_dataset.py- Custom dataset implementationexamples/custom_model.py- Custom model implementationexamples/augmentation_example.py- Data augmentation exampleexamples/config_example.yaml- Configuration file exampleexamples/pipeline_example.sh- Pipeline script example