releso.agent.PPOAgent
- class releso.agent.PPOAgent(*, save_location: Path, logger_name: str | None = None, tensorboard_log: str | None = None, policy: Literal['MlpPolicy', 'CnnPolicy', 'MultiInputPolicy'], use_custom_feature_extractor: Literal['resnet18', 'mobilenet_v2'] | None = None, cfe_without_linear: bool = False, policy_kwargs: Dict[str, Any] | None = None, type: Literal['PPO'], learning_rate: float = 0.0003, n_steps: int = 2048, batch_size: int | None = 64, n_epochs: int = 10, gamma: float = 0.99, gae_lambda: float = 0.95, clip_range: float = 0.2, ent_coef: float = 0.0, vf_coef: float = 0.5, seed: int | None = None, device: str = 'auto')
Bases:
BaseTrainingAgentPPO agent definition.
PPO definition for the stable_baselines3 implementation for this algorithm. Variable comments are taken from the stable_baselines3 documentation.
- __init__(**data: Any) None
Constructor for the ReLeSO basemodel object.
Methods
convert_to_pathlib_add_datetime(v)Add timestamp to save_location, of applicable.
get_additional_kwargs(**kwargs)Add additional keyword arguments for agent instantiation.
get_agent(environment[, normalizer_divisor])Creates the stable_baselines version of the wanted Agent.
get_logger()Gets the currently defined environment logger.
get_next_tensorboard_experiment_name()Return tensorboard experiment name.
set_logger_name_recursively(logger_name)Set the logger_name variable for all child elements.
Attributes
What RL algorithm was used to train the agent.
The learning rate, it can be a function of the current progress remaining (from 1 to 0)
The number of steps to run for each environment per update(i.e. rollout buffer size is n_steps * n_envs where n_envs is number of environment copies running in parallel) NOTE: n_steps * n_envs must be greater than 1 (because of the advantage normalization) See https://github.com/pytorch/pytorch/issues/29372.
Minibatch size
Number of epoch when optimizing the surrogate loss
Discount factor
Factor for trade-off of bias vs variance for Generalized Advantage Estimator
Clipping parameter, it can be a function of the current progress remaining (from 1 to 0).
Entropy coefficient for the loss calculation
Value function coefficient for the loss calculation
Seed for the pseudo random generators
Device (cpu, cuda, …) on which the code should be run.
policypolicy defines the network structure which the agent uses
use_custom_feature_extractorIf given the str identifies the Custom Feature Extractor to be added.
cfe_without_linearuse the custom feature extractor with out a final linear layer
policy_kwargsadditional arguments to be passed to the policy on creation
tensorboard_logbase directory of the tensorboard logs if given an experiment name with a current timestamp is also added.
save_locationDefinition of the save location of the logs and validation results.
logger_namename of the logger.
- agent_type: Literal['PPO']
What RL algorithm was used to train the agent. Needs to be know to correctly load the agent.
- batch_size: int | None
Minibatch size
- clip_range: float
Clipping parameter, it can be a function of the current progress remaining (from 1 to 0).
- device: str
Device (cpu, cuda, …) on which the code should be run. Setting it to auto, the code will be run on the GPU if possible.
- ent_coef: float
Entropy coefficient for the loss calculation
- gae_lambda: float
Factor for trade-off of bias vs variance for Generalized Advantage Estimator
- gamma: float
Discount factor
- get_agent(environment: GymEnvironment, normalizer_divisor: int = 1) PPO
Creates the stable_baselines version of the wanted Agent.
Uses all Variables given in the object (except agent_type) as the input parameters of the agent object creation.
Notes
The PPO variable n_steps scales with the amount of environments the agent is trained with. If this behavior is unwanted you can set the variable ‘’normalize_training_values’’ in BaseParser to true. This will set this variable to the number of environments so that the scaled values can be unscaled.
- Parameters:
environment (GymEnvironment) – The environment the agent uses to train.
normalizer_divisor (int) – Divides the variable n_steps by this value to descale scaled values with n_environments unequal zero.
- Returns:
Initialized PPO agent.
- Return type:
PPO
- learning_rate: float
The learning rate, it can be a function of the current progress remaining (from 1 to 0)
- n_epochs: int
Number of epoch when optimizing the surrogate loss
- n_steps: int
The number of steps to run for each environment per update(i.e. rollout buffer size is n_steps * n_envs where n_envs is number of environment copies running in parallel) NOTE: n_steps * n_envs must be greater than 1 (because of the advantage normalization) See https://github.com/pytorch/pytorch/issues/29372
- seed: int | None
Seed for the pseudo random generators
- vf_coef: float
Value function coefficient for the loss calculation