Class: Google::Apis::AiplatformV1beta1::GoogleCloudAiplatformV1beta1ValidateReinforcementTuningRewardRequest

Inherits:
Object
  • Object
show all
Includes:
Core::Hashable, Core::JsonObjectSupport
Defined in:
lib/google/apis/aiplatform_v1beta1/classes.rb,
lib/google/apis/aiplatform_v1beta1/representations.rb,
lib/google/apis/aiplatform_v1beta1/representations.rb

Overview

Request message for GenAiTuningService.ValidateReinforcementTuningReward.

Instance Attribute Summary collapse

Instance Method Summary collapse

Constructor Details

#initialize(**args) ⇒ GoogleCloudAiplatformV1beta1ValidateReinforcementTuningRewardRequest

Returns a new instance of GoogleCloudAiplatformV1beta1ValidateReinforcementTuningRewardRequest.



68798
68799
68800
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 68798

def initialize(**args)
   update!(**args)
end

Instance Attribute Details

#composite_reward_configGoogle::Apis::AiplatformV1beta1::GoogleCloudAiplatformV1beta1CompositeReinforcementTuningRewardConfig

Composite reward function configuration for reinforcement tuning. Corresponds to the JSON property compositeRewardConfig



68775
68776
68777
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 68775

def composite_reward_config
  @composite_reward_config
end

#exampleGoogle::Apis::AiplatformV1beta1::GoogleCloudAiplatformV1beta1ReinforcementTuningExample

User-facing format for Gemini Reinforcement Tuning examples on Vertex. Corresponds to the JSON property example



68780
68781
68782
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 68780

def example
  @example
end

#sample_responseGoogle::Apis::AiplatformV1beta1::GoogleCloudAiplatformV1beta1Content

The structured data content of a message. A Content message contains a role field, which indicates the producer of the content, and a parts field, which contains the multi-part data of the message. Corresponds to the JSON property sampleResponse



68787
68788
68789
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 68787

def sample_response
  @sample_response
end

#single_reward_configGoogle::Apis::AiplatformV1beta1::GoogleCloudAiplatformV1beta1SingleReinforcementTuningRewardConfig

SingleReinforcementTuningRewardConfig defines a single reward function configuration for RL tuning. Each reward calculation/evaluation consists of two stages: 1. Stage 1: Parses the part of information important from sample response via regex extract, or simply takes the sample response unmodified. 2. Stage 2: Calls the configured reward scorer to compute the reward. Corresponds to the JSON property singleRewardConfig



68796
68797
68798
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 68796

def single_reward_config
  @single_reward_config
end

Instance Method Details

#update!(**args) ⇒ Object

Update properties of this object



68803
68804
68805
68806
68807
68808
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 68803

def update!(**args)
  @composite_reward_config = args[:composite_reward_config] if args.key?(:composite_reward_config)
  @example = args[:example] if args.key?(:example)
  @sample_response = args[:sample_response] if args.key?(:sample_response)
  @single_reward_config = args[:single_reward_config] if args.key?(:single_reward_config)
end