Module: Protocol::URL::Encoding
- Defined in:
- lib/protocol/url/encoding.rb
Overview
Helpers for encoding and decoding URL components.
Constant Summary collapse
- NON_PATH_CHARACTER_PATTERN =
Matches characters that are not allowed in a URI path segment. According to RFC 3986 Section 3.3 (https://tools.ietf.org/html/rfc3986#section-3.3), a valid path segment consists of "pchar" characters. This pattern identifies characters that must be percent-encoded when included in a URI path segment.
/([^a-zA-Z0-9_\-\.~!$&'()*+,;=:@\/]+)/.freeze
- NON_FRAGMENT_CHARACTER_PATTERN =
Matches characters that are not allowed in a URI fragment. According to RFC 3986 Section 3.5, a valid fragment consists of pchar / "/" / "?" characters.
/([^a-zA-Z0-9_\-\.~!$&'()*+,;=:@\/\?]+)/.freeze
Class Method Summary collapse
-
.assign(keys, value, parent) ⇒ Object
Assign a value to a nested hash.
-
.decode(string, maximum = 8, symbolize_keys: false) ⇒ Object
Decode a URL-encoded query string into a hash.
-
.decode_www_form(string, maximum = 8, symbolize_keys: false) ⇒ Object
Decode an
application/x-www-form-urlencodedstring into a hash. -
.encode(value, prefix = nil) ⇒ Object
Encodes a hash or array into a query string.
-
.escape(string, encoding = string.encoding) ⇒ Object
Escapes a string using percent encoding, e.g.
-
.escape_fragment(fragment) ⇒ Object
Escapes non-fragment characters using percent encoding.
-
.escape_path(path) ⇒ Object
Escapes non-path characters using percent encoding.
-
.scan(string) ⇒ Object
Scan a string for URL-encoded key/value pairs.
-
.split(name) ⇒ Object
Split a key into parts, e.g.
-
.unescape(string, encoding = string.encoding) ⇒ Object
Unescapes a percent encoded string, e.g.
-
.unescape_path(string, encoding = string.encoding) ⇒ Object
Unescapes a percent encoded path component, preserving encoded path separators.
Class Method Details
.assign(keys, value, parent) ⇒ Object
Assign a value to a nested hash.
This method handles building nested data structures from query string parameters, including arrays of objects. When processing array elements (empty key like []), it intelligently decides whether to add to the last array element or create a new one.
174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 |
# File 'lib/protocol/url/encoding.rb', line 174 def self.assign(keys, value, parent) top, *middle = keys middle.each_with_index do |key, index| if key.nil? or key.empty? # Array element (e.g., items[]): parent = (parent[top] ||= Array.new) top = parent.size # Check if we should reuse the last array element or create a new one. If there's a nested key coming next, and the last array element already has that key, then we need a new array element. Otherwise, add to the existing one. if nested = middle[index+1] and last = parent.last # If the last element doesn't include the nested key, reuse it (decrement index). # If it does include the key, keep current index (creates new element). top -= 1 unless last.include?(nested) end else # Hash key (e.g., user[name]): parent = (parent[top] ||= Hash.new) top = key end end parent[top] = value end |
.decode(string, maximum = 8, symbolize_keys: false) ⇒ Object
Decode a URL-encoded query string into a hash.
213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 |
# File 'lib/protocol/url/encoding.rb', line 213 def self.decode(string, maximum = 8, symbolize_keys: false) parameters = {} self.scan(string) do |name, value| keys = self.split(name) if keys.empty? raise ArgumentError, "Invalid key path: #{name.inspect}!" end if keys.size > maximum raise ArgumentError, "Key length exceeded limit!" end if symbolize_keys keys.collect!{|key| key.empty? ? nil : key.to_sym} end self.assign(keys, value, parameters) end return parameters end |
.decode_www_form(string, maximum = 8, symbolize_keys: false) ⇒ Object
Decode an application/x-www-form-urlencoded string into a hash.
In addition to percent encoding, this format represents spaces using +.
244 245 246 |
# File 'lib/protocol/url/encoding.rb', line 244 def self.decode_www_form(string, maximum = 8, symbolize_keys: false) return self.decode(string.gsub("+", "%20"), maximum, symbolize_keys: symbolize_keys) end |
.encode(value, prefix = nil) ⇒ Object
Encodes a hash or array into a query string. This method is used to encode query parameters in a URL. For example, {"a" => 1, "b" => 2} is encoded as a=1&b=2.
118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 |
# File 'lib/protocol/url/encoding.rb', line 118 def self.encode(value, prefix = nil) case value when Array return value.map {|v| self.encode(v, "#{prefix}[]") }.join("&") when Hash return value.map {|k, v| self.encode(v, prefix ? "#{prefix}[#{escape(k.to_s)}]" : escape(k.to_s)) }.reject(&:empty?).join("&") when nil return prefix else raise ArgumentError, "value must be a Hash" if prefix.nil? return "#{prefix}=#{escape(value.to_s)}" end end |
.escape(string, encoding = string.encoding) ⇒ Object
Escapes a string using percent encoding, e.g. a b -> a%20b.
22 23 24 25 26 |
# File 'lib/protocol/url/encoding.rb', line 22 def self.escape(string, encoding = string.encoding) string.b.gsub(/([^a-zA-Z0-9_.\-]+)/) do |m| "%" + m.unpack("H2" * m.bytesize).join("%").upcase end.force_encoding(encoding) end |
.escape_fragment(fragment) ⇒ Object
Escapes non-fragment characters using percent encoding. According to RFC 3986 Section 3.5, fragments can contain pchar / "/" / "?" characters.
99 100 101 102 103 104 |
# File 'lib/protocol/url/encoding.rb', line 99 def self.escape_fragment(fragment) encoding = fragment.encoding fragment.b.gsub(NON_FRAGMENT_CHARACTER_PATTERN) do |m| "%" + m.unpack("H2" * m.bytesize).join("%").upcase end.force_encoding(encoding) end |
.escape_path(path) ⇒ Object
Escapes non-path characters using percent encoding. In other words, this method escapes characters that are not allowed in a URI path segment. According to RFC 3986 Section 3.3 (https://tools.ietf.org/html/rfc3986#section-3.3), a valid path segment consists of "pchar" characters. This method percent-encodes characters that are not "pchar" characters.
88 89 90 91 92 93 |
# File 'lib/protocol/url/encoding.rb', line 88 def self.escape_path(path) encoding = path.encoding path.b.gsub(NON_PATH_CHARACTER_PATTERN) do |m| "%" + m.unpack("H2" * m.bytesize).join("%").upcase end.force_encoding(encoding) end |
.scan(string) ⇒ Object
Scan a string for URL-encoded key/value pairs.
141 142 143 144 145 146 147 148 149 |
# File 'lib/protocol/url/encoding.rb', line 141 def self.scan(string) string.split("&") do |assignment| next if assignment.empty? key, value = assignment.split("=", 2) yield unescape(key), value.nil? ? value : unescape(value) end end |
.split(name) ⇒ Object
Split a key into parts, e.g. a[b][c] -> ["a", "b", "c"].
155 156 157 158 159 160 |
# File 'lib/protocol/url/encoding.rb', line 155 def self.split(name) name.scan(/([^\[]+)|(?:\[(.*?)\])/)&.tap do |parts| parts.flatten! parts.compact! end end |
.unescape(string, encoding = string.encoding) ⇒ Object
Unescapes a percent encoded string, e.g. a%20b -> a b.
40 41 42 43 44 |
# File 'lib/protocol/url/encoding.rb', line 40 def self.unescape(string, encoding = string.encoding) string.b.gsub(/%(\h\h)/) do |hex| Integer($1, 16).chr end.force_encoding(encoding) end |
.unescape_path(string, encoding = string.encoding) ⇒ Object
Unescapes a percent encoded path component, preserving encoded path separators.
This method unescapes percent-encoded characters except for path separators
(forward slash / and backslash \). This prevents encoded separators like
%2F or %5C from being decoded into actual path separators, which could
allow bypassing path component boundaries.
60 61 62 63 64 65 66 67 68 69 70 71 72 |
# File 'lib/protocol/url/encoding.rb', line 60 def self.unescape_path(string, encoding = string.encoding) string.b.gsub(/%(\h\h)/) do |hex| byte = Integer($1, 16) char = byte.chr # Don't decode forward slash (0x2F) or backslash (0x5C) if byte == 0x2F || byte == 0x5C hex # Keep as %2F or %5C else char end end.force_encoding(encoding) end |