Class: Google::Apis::AiplatformV1beta1::GoogleCloudAiplatformV1beta1ValidateReinforcementTuningRewardRequest

Inherits:
Object
  • Object
show all
Includes:
Core::Hashable, Core::JsonObjectSupport
Defined in:
lib/google/apis/aiplatform_v1beta1/classes.rb,
lib/google/apis/aiplatform_v1beta1/representations.rb,
lib/google/apis/aiplatform_v1beta1/representations.rb

Overview

Request message for GenAiTuningService.ValidateReinforcementTuningReward.

Instance Attribute Summary collapse

Instance Method Summary collapse

Constructor Details

#initialize(**args) ⇒ GoogleCloudAiplatformV1beta1ValidateReinforcementTuningRewardRequest

Returns a new instance of GoogleCloudAiplatformV1beta1ValidateReinforcementTuningRewardRequest.



67970
67971
67972
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 67970

def initialize(**args)
   update!(**args)
end

Instance Attribute Details

#composite_reward_configGoogle::Apis::AiplatformV1beta1::GoogleCloudAiplatformV1beta1CompositeReinforcementTuningRewardConfig

Composite reward function configuration for reinforcement tuning. Corresponds to the JSON property compositeRewardConfig



67947
67948
67949
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 67947

def composite_reward_config
  @composite_reward_config
end

#exampleGoogle::Apis::AiplatformV1beta1::GoogleCloudAiplatformV1beta1ReinforcementTuningExample

User-facing format for Gemini Reinforcement Tuning examples on Vertex. Corresponds to the JSON property example



67952
67953
67954
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 67952

def example
  @example
end

#sample_responseGoogle::Apis::AiplatformV1beta1::GoogleCloudAiplatformV1beta1Content

The structured data content of a message. A Content message contains a role field, which indicates the producer of the content, and a parts field, which contains the multi-part data of the message. Corresponds to the JSON property sampleResponse



67959
67960
67961
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 67959

def sample_response
  @sample_response
end

#single_reward_configGoogle::Apis::AiplatformV1beta1::GoogleCloudAiplatformV1beta1SingleReinforcementTuningRewardConfig

SingleReinforcementTuningRewardConfig defines a single reward function configuration for RL tuning. Each reward calculation/evaluation consists of two stages: 1. Stage 1: Parses the part of information important from sample response via regex extract, or simply takes the sample response unmodified. 2. Stage 2: Calls the configured reward scorer to compute the reward. Corresponds to the JSON property singleRewardConfig



67968
67969
67970
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 67968

def single_reward_config
  @single_reward_config
end

Instance Method Details

#update!(**args) ⇒ Object

Update properties of this object



67975
67976
67977
67978
67979
67980
# File 'lib/google/apis/aiplatform_v1beta1/classes.rb', line 67975

def update!(**args)
  @composite_reward_config = args[:composite_reward_config] if args.key?(:composite_reward_config)
  @example = args[:example] if args.key?(:example)
  @sample_response = args[:sample_response] if args.key?(:sample_response)
  @single_reward_config = args[:single_reward_config] if args.key?(:single_reward_config)
end