โมเดล Seq2seq (ลำดับต่อลำดับ) ด้วย PyTorch

⚡ สรุปอย่างชาญฉลาด

Seq2Seq เป็นสถาปัตยกรรมตัวเข้ารหัส-ตัวถอดรหัสที่แปลงลำดับอินพุตเป็นลำดับเอาต์พุตโดยใช้โครงข่ายประสาทเทียมแบบวนซ้ำสองโครงข่าย ซึ่งเป็นหัวใจสำคัญของการแปลภาษาด้วยเครื่องจักรและงานประมวลผลภาษาธรรมชาติอื่นๆ ที่ความยาวของอินพุตและเอาต์พุตแตกต่างกัน

  • 🧠 พื้นฐานของ NLP: การประมวลผลภาษาธรรมชาติช่วยให้คอมพิวเตอร์เข้าใจและตอบสนองต่อภาษามนุษย์ได้ ดังที่เห็นได้ใน Google Translate.
  • 🔄 ตัวเข้ารหัส-ตัวถอดรหัส: RNN ตัวแรกทำหน้าที่เข้ารหัสข้อมูลนำเข้าให้เป็นสถานะ และ RNN ตัวที่สองทำหน้าที่ถอดรหัสสถานะนั้นให้เป็นข้อมูลส่งออก
  • 🧮 ชั้นต่างๆ ของ GRU: หน่วยที่เกิดซ้ำ Gated track สถานะที่ซ่อนอยู่และการรีเซ็ตการอัปเดต การอัปเดต และเกตใหม่ตลอดลำดับ
  • 🪙 ราชสกุล: เอสโอเอสและ EOS โทเค็นทำหน้าที่ระบุจุดเริ่มต้นและจุดสิ้นสุดของแต่ละลำดับในระหว่างการฝึกฝนและการทำนาย
  • 🎯 ครูบังคับ: การป้อนคำจริงแทนคำที่คาดเดาจะช่วยให้การฝึกฝนมีเสถียรภาพและรวดเร็วยิ่งขึ้น
  • 🤖 ผลกระทบของ AI: Seq2Seq เป็นพื้นฐานของการแปล การสรุปความ และแชทบอท และเป็นแรงบันดาลใจให้กับโมเดล Attention และ Transformer ในเวลาต่อมา

แบบจำลองลำดับต่อลำดับ Seq2seq

NLP คืออะไร

NLP หรือการประมวลผลภาษาธรรมชาติ เป็นหนึ่งในสาขาที่ได้รับความนิยมของปัญญาประดิษฐ์ ซึ่งช่วยให้คอมพิวเตอร์เข้าใจ จัดการ หรือตอบสนองต่อมนุษย์ด้วยภาษาธรรมชาติของพวกเขา NLP คือกลไกสำคัญเบื้องหลัง Google Translate ที่ช่วยให้เราเข้าใจภาษาอื่นๆ

Seq2Seq คืออะไร?

Seq2Seq เป็นวิธีการแปลด้วยเครื่องโดยใช้ตัวเข้ารหัสและตัวถอดรหัสซึ่งแมปอินพุตของลำดับกับเอาต์พุตของลำดับด้วยแท็กและค่าความสนใจ แนวคิดคือการใช้ 2 RNN ที่จะทำงานร่วมกับโทเค็นพิเศษและพยายามทำนายลำดับสถานะถัดไปจากลำดับก่อนหน้า

วิธีทำนายลำดับจากลำดับก่อนหน้า

ทำนายลำดับจากลำดับก่อนหน้า

ต่อไปนี้คือขั้นตอนในการทำนายลำดับจากลำดับก่อนหน้าโดยใช้ PyTorCH

ขั้นตอนที่ 1) กำลังโหลดข้อมูลของเรา

สำหรับชุดข้อมูลของเรา คุณจะใช้ชุดข้อมูลจาก คู่ประโยคสองภาษาที่คั่นด้วยแท็บ- ที่นี่ฉันจะใช้ชุดข้อมูลภาษาอังกฤษเป็นภาษาอินโดนีเซีย คุณสามารถเลือกสิ่งที่คุณต้องการได้ แต่อย่าลืมเปลี่ยนชื่อไฟล์และไดเร็กทอรีในโค้ด

from __future__ import unicode_literals, print_function, division
import torch
import torch.nn as nn
import torch.optim as optim
import torch.nn.functional as F

import numpy as np
import pandas as pd

import os
import re
import random

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

ขั้นตอนที่ 2) การเตรียมข้อมูล

คุณไม่สามารถใช้ชุดข้อมูลโดยตรงได้ คุณต้องแยกประโยคออกเป็นคำๆ แล้วแปลงเป็นเวกเตอร์แบบ One-Hot ก่อน จากนั้นแต่ละคำจะถูกกำหนดดัชนีที่ไม่ซ้ำกันในคลาส Lang เพื่อสร้างพจนานุกรม คลาส Lang จะจัดเก็บแต่ละประโยคและแยกคำต่อคำด้วยเมธอด addSentence จากนั้นสร้างพจนานุกรมโดยการกำหนดดัชนีให้กับคำที่ไม่รู้จักแต่ละคำสำหรับโมเดลแบบลำดับต่อลำดับ (Sequence to sequence models)

SOS_token = 0
EOS_token = 1
MAX_LENGTH = 20

#initialize Lang Class
class Lang:
   def __init__(self):
       #initialize containers to hold the words and corresponding index
       self.word2index = {}
       self.word2count = {}
       self.index2word = {0: "SOS", 1: "EOS"}
       self.n_words = 2  # Count SOS and EOS

#split a sentence into words and add it to the container
   def addSentence(self, sentence):
       for word in sentence.split(' '):
           self.addWord(word)

#If the word is not in the container, the word will be added to it, else, update the word counter
   def addWord(self, word):
       if word not in self.word2index:
           self.word2index[word] = self.n_words
           self.word2count[word] = 1
           self.index2word[self.n_words] = word
           self.n_words += 1
       else:
           self.word2count[word] += 1

คลาส Lang เป็นคลาสที่จะช่วยเราสร้างพจนานุกรม สำหรับแต่ละภาษา ประโยคทุกประโยคจะถูกแยกออกเป็นคำๆ แล้วเพิ่มเข้าไปในคอนเทนเนอร์ แต่ละคอนเทนเนอร์จะเก็บคำต่างๆ ไว้ในดัชนีที่เหมาะสม นับจำนวนคำ และเพิ่มดัชนีของคำนั้น เพื่อให้เราสามารถใช้ดัชนีนั้นในการค้นหาดัชนีของคำ หรือค้นหาคำจากดัชนีของมันได้

เนื่องจากข้อมูลของเราถูกคั่นด้วย TAB คุณจึงจำเป็นต้องใช้ หมีแพนด้า ในฐานะตัวโหลดข้อมูลของเรา Pandas จะอ่านข้อมูลของเราในรูปแบบ DataFrame และแยกออกเป็นประโยคต้นฉบับและประโยคเป้าหมาย สำหรับแต่ละประโยคที่คุณมี คุณจะต้องแปลงให้เป็นตัวพิมพ์เล็ก ลบอักขระที่ไม่ใช่ตัวพิมพ์ใหญ่ แปลงจาก Unicode เป็น ASCII และแยกประโยคเพื่อให้ได้แต่ละคำออกมา

#Normalize every sentence
def normalize_sentence(df, lang):
   sentence = df[lang].str.lower()
   sentence = sentence.str.replace('[^A-Za-z\s]+', '')
   sentence = sentence.str.normalize('NFD')
   sentence = sentence.str.encode('ascii', errors='ignore').str.decode('utf-8')
   return sentence

def read_sentence(df, lang1, lang2):
   sentence1 = normalize_sentence(df, lang1)
   sentence2 = normalize_sentence(df, lang2)
   return sentence1, sentence2

def read_file(loc, lang1, lang2):
   df = pd.read_csv(loc, delimiter='\t', header=None, names=[lang1, lang2])
   return df

def process_data(lang1,lang2):
   df = read_file('text/%s-%s.txt' % (lang1, lang2), lang1, lang2)
   print("Read %s sentence pairs" % len(df))
   sentence1, sentence2 = read_sentence(df, lang1, lang2)

   source = Lang()
   target = Lang()
   pairs = []
   for i in range(len(df)):
       if len(sentence1[i].split(' ')) < MAX_LENGTH and len(sentence2[i].split(' ')) < MAX_LENGTH:
           full = [sentence1[i], sentence2[i]]
           source.addSentence(sentence1[i])
           target.addSentence(sentence2[i])
           pairs.append(full)

   return source, target, pairs

ฟังก์ชันที่มีประโยชน์อีกอย่างหนึ่งที่คุณจะได้ใช้คือการแปลงคู่ข้อมูลเป็นเทนเซอร์ นี่เป็นสิ่งสำคัญมากเพราะเครือข่ายของเราอ่านได้เฉพาะข้อมูลประเภทเทนเซอร์เท่านั้น นอกจากนี้ยังสำคัญเพราะนี่คือส่วนที่ในแต่ละตอนของประโยคจะมีโทเค็นเพื่อบอกเครือข่ายว่าข้อมูลป้อนเข้าสิ้นสุดลงแล้ว สำหรับแต่ละคำในประโยค มันจะดึงดัชนีจากคำที่เหมาะสมในพจนานุกรมและเพิ่มโทเค็นไว้ที่ท้ายประโยค

def indexesFromSentence(lang, sentence):
   return [lang.word2index[word] for word in sentence.split(' ')]

def tensorFromSentence(lang, sentence):
   indexes = indexesFromSentence(lang, sentence)
   indexes.append(EOS_token)
   return torch.tensor(indexes, dtype=torch.long, device=device).view(-1, 1)

def tensorsFromPair(input_lang, output_lang, pair):
   input_tensor = tensorFromSentence(input_lang, pair[0])
   target_tensor = tensorFromSentence(output_lang, pair[1])
   return (input_tensor, target_tensor)

โมเดล Seq2Seq

โมเดล Seq2seq

PyTorโมเดล ch Seq2seq เป็นโมเดลประเภทหนึ่งที่ใช้ PyTorโมเดลประกอบด้วยตัวเข้ารหัสและตัวถอดรหัส ตัวเข้ารหัสจะเข้ารหัสคำในประโยคทีละคำลงในดัชนีคำศัพท์หรือคำที่รู้จัก ส่วนตัวถอดรหัสจะทำนายผลลัพธ์ของข้อมูลที่เข้ารหัสโดยการถอดรหัสข้อมูลตามลำดับ และจะพยายามใช้ข้อมูลล่าสุดเป็นข้อมูลถัดไปหากเป็นไปได้ ด้วยวิธีนี้ ยังสามารถทำนายข้อมูลถัดไปเพื่อสร้างประโยคได้อีกด้วย แต่ละประโยคจะได้รับโทเค็นเพื่อทำเครื่องหมายจุดสิ้นสุดของลำดับ เมื่อสิ้นสุดการทำนาย จะมีโทเค็นเพื่อทำเครื่องหมายจุดสิ้นสุดของผลลัพธ์เช่นกัน ดังนั้น จากตัวเข้ารหัส จะส่งสถานะไปยังตัวถอดรหัสเพื่อทำนายผลลัพธ์

โมเดล Seq2seq

ตัวเข้ารหัส (Encoder) จะเข้ารหัสประโยคอินพุตของเราทีละคำตามลำดับ และในตอนท้ายจะมีโทเค็นเพื่อระบุจุดสิ้นสุดของประโยค ตัวเข้ารหัสประกอบด้วยเลเยอร์ฝังตัว (Embedding layer) และเลเยอร์ GRU เลเยอร์ฝังตัวเป็นตารางค้นหาที่เก็บค่าฝังตัวของอินพุตของเราลงในพจนานุกรมคำที่มีขนาดคงที่ ซึ่งจะถูกส่งต่อไปยังเลเยอร์ GRU เลเยอร์ GRU เป็นหน่วยวนซ้ำแบบมีเกต (Gated Recurrent Unit) ที่ประกอบด้วยโครงสร้างหลายชั้น ร.น. ที่จะคำนวณอินพุตตามลำดับ เลเยอร์นี้จะคำนวณสถานะที่ซ่อนอยู่จากสถานะก่อนหน้า และอัปเดตการรีเซ็ต อัปเดต และเกตใหม่

โมเดล Seq2seq

ตัวถอดรหัส (Decoder) จะถอดรหัสข้อมูลอินพุตจากเอาต์พุตของตัวเข้ารหัส (Encoder) โดยจะพยายามทำนายเอาต์พุตถัดไปและพยายามใช้เป็นอินพุตถัดไปหากเป็นไปได้ ตัวถอดรหัสประกอบด้วยเลเยอร์ฝังตัว (Embedding layer), เลเยอร์ GRU (Global Restricted Resource) และเลเยอร์เชิงเส้น (Linear layer) เลเยอร์ฝังตัวจะสร้างตารางค้นหาสำหรับเอาต์พุตและส่งต่อไปยังเลเยอร์ GRU เพื่อคำนวณสถานะเอาต์พุตที่ทำนายได้ หลังจากนั้น เลเยอร์เชิงเส้นจะช่วยคำนวณฟังก์ชันการกระตุ้น (Activation function) เพื่อกำหนดค่าที่แท้จริงของเอาต์พุตที่ทำนายได้

class Encoder(nn.Module):
   def __init__(self, input_dim, hidden_dim, embbed_dim, num_layers):
       super(Encoder, self).__init__()

       #set the encoder input dimesion , embbed dimesion, hidden dimesion, and number of layers
       self.input_dim = input_dim
       self.embbed_dim = embbed_dim
       self.hidden_dim = hidden_dim
       self.num_layers = num_layers

       #initialize the embedding layer with input and embbed dimention
       self.embedding = nn.Embedding(input_dim, self.embbed_dim)
       #intialize the GRU to take the input dimetion of embbed, and output dimention of hidden and
       #set the number of gru layers
       self.gru = nn.GRU(self.embbed_dim, self.hidden_dim, num_layers=self.num_layers)

   def forward(self, src):
       embedded = self.embedding(src).view(1,1,-1)
       outputs, hidden = self.gru(embedded)
       return outputs, hidden

class Decoder(nn.Module):
   def __init__(self, output_dim, hidden_dim, embbed_dim, num_layers):
       super(Decoder, self).__init__()

#set the encoder output dimension, embed dimension, hidden dimension, and number of layers
       self.embbed_dim = embbed_dim
       self.hidden_dim = hidden_dim
       self.output_dim = output_dim
       self.num_layers = num_layers

# initialize every layer with the appropriate dimension. For the decoder layer, it will consist of an embedding, GRU, a Linear layer and a Log softmax activation function.
       self.embedding = nn.Embedding(output_dim, self.embbed_dim)
       self.gru = nn.GRU(self.embbed_dim, self.hidden_dim, num_layers=self.num_layers)
       self.out = nn.Linear(self.hidden_dim, output_dim)
       self.softmax = nn.LogSoftmax(dim=1)

   def forward(self, input, hidden):
# reshape the input to (1, batch_size)
       input = input.view(1, -1)
       embedded = F.relu(self.embedding(input))
       output, hidden = self.gru(embedded, hidden)
       prediction = self.softmax(self.out(output[0]))
       return prediction, hidden

class Seq2Seq(nn.Module):
   def __init__(self, encoder, decoder, device, MAX_LENGTH=MAX_LENGTH):
       super().__init__()

#initialize the encoder and decoder
       self.encoder = encoder
       self.decoder = decoder
       self.device = device

   def forward(self, source, target, teacher_forcing_ratio=0.5):
       input_length = source.size(0) #get the input length (number of words in sentence)
       batch_size = target.shape[1]
       target_length = target.shape[0]
       vocab_size = self.decoder.output_dim

#initialize a variable to hold the predicted outputs
       outputs = torch.zeros(target_length, batch_size, vocab_size).to(self.device)

#encode every word in a sentence
       for i in range(input_length):
           encoder_output, encoder_hidden = self.encoder(source[i])

#use the encoder's hidden layer as the decoder hidden
       decoder_hidden = encoder_hidden.to(device)

#add a token before the first predicted word
       decoder_input = torch.tensor([SOS_token], device=device)  # SOS

#topk is used to get the top K value over a list
#predict the output word from the current target word. If we enable the teaching force, then the next decoder input is the next word, else, use the decoder output highest value.
       for t in range(target_length):
           decoder_output, decoder_hidden = self.decoder(decoder_input, decoder_hidden)
           outputs[t] = decoder_output
           teacher_force = random.random() < teacher_forcing_ratio
           topv, topi = decoder_output.topk(1)
           input = (target[t] if teacher_force else topi)
           if(teacher_force == False and input.item() == EOS_token):
               break

       return outputs

ขั้นตอนที่ 3) การฝึกโมเดล

กระบวนการฝึกฝนในโมเดล Seq2seq เริ่มต้นด้วยการแปลงประโยคแต่ละคู่ให้เป็น Tensor จากดัชนี Lang โมเดลลำดับต่อลำดับของเราจะใช้ SGD เป็นตัวปรับแต่ง และฟังก์ชัน NLLLoss เพื่อคำนวณค่าความสูญเสีย กระบวนการฝึกฝนเริ่มต้นด้วยการป้อนประโยคแต่ละคู่ให้กับโมเดลเพื่อทำนายผลลัพธ์ที่ถูกต้อง ในแต่ละขั้นตอน ผลลัพธ์จากโมเดลจะถูกคำนวณร่วมกับคำจริงเพื่อหาค่าความสูญเสียและอัปเดตพารามิเตอร์ ดังนั้น เนื่องจากคุณจะใช้การวนซ้ำ 75000 ครั้ง โมเดลลำดับต่อลำดับของเราจะสร้างคู่แบบสุ่ม 75000 คู่จากชุดข้อมูลของเรา

teacher_forcing_ratio = 0.5

def clacModel(model, input_tensor, target_tensor, model_optimizer, criterion):
   model_optimizer.zero_grad()

   input_length = input_tensor.size(0)
   loss = 0
   epoch_loss = 0
   # print(input_tensor.shape)

   output = model(input_tensor, target_tensor)

   num_iter = output.size(0)
   print(num_iter)

#calculate the loss from a predicted sentence with the expected result
   for ot in range(num_iter):
       loss += criterion(output[ot], target_tensor[ot])

   loss.backward()
   model_optimizer.step()
   epoch_loss = loss.item() / num_iter

   return epoch_loss

def trainModel(model, source, target, pairs, num_iteration=20000):
   model.train()

   optimizer = optim.SGD(model.parameters(), lr=0.01)
   criterion = nn.NLLLoss()
   total_loss_iterations = 0

   training_pairs = [tensorsFromPair(source, target, random.choice(pairs))
                     for i in range(num_iteration)]

   for iter in range(1, num_iteration+1):
       training_pair = training_pairs[iter - 1]
       input_tensor = training_pair[0]
       target_tensor = training_pair[1]

       loss = clacModel(model, input_tensor, target_tensor, optimizer, criterion)

       total_loss_iterations += loss

       if iter % 5000 == 0:
           avarage_loss= total_loss_iterations / 5000
           total_loss_iterations = 0
           print('%d %.4f' % (iter, avarage_loss))

   torch.save(model.state_dict(), 'mytraining.pt')
   return model

ขั้นตอนที่ 4) ทดสอบโมเดล

กระบวนการประเมินของ Seq2seq PyTorch คือการตรวจสอบผลลัพธ์ของโมเดล แต่ละคู่ของโมเดลแบบลำดับต่อลำดับจะถูกป้อนเข้าไปในโมเดลและสร้างคำที่คาดการณ์ได้ หลังจากนั้น คุณจะดูค่าสูงสุดในแต่ละผลลัพธ์เพื่อค้นหาดัชนีที่ถูกต้อง และในตอนท้าย คุณจะเปรียบเทียบเพื่อดูว่าการคาดการณ์ของโมเดลของเราตรงกับประโยคจริงหรือไม่

def evaluate(model, input_lang, output_lang, sentences, max_length=MAX_LENGTH):
   with torch.no_grad():
       input_tensor = tensorFromSentence(input_lang, sentences[0])
       output_tensor = tensorFromSentence(output_lang, sentences[1])

       decoded_words = []

       output = model(input_tensor, output_tensor)
       # print(output_tensor)

       for ot in range(output.size(0)):
           topv, topi = output[ot].topk(1)
           # print(topi)

           if topi[0].item() == EOS_token:
               decoded_words.append('')
               break
           else:
               decoded_words.append(output_lang.index2word[topi[0].item()])
   return decoded_words

def evaluateRandomly(model, source, target, pairs, n=10):
   for i in range(n):
       pair = random.choice(pairs)
       print('source {}'.format(pair[0]))
       print('target {}'.format(pair[1]))
       output_words = evaluate(model, source, target, pair)
       output_sentence = ' '.join(output_words)
       print('predicted {}'.format(output_sentence))

ตอนนี้ เรามาเริ่มการฝึกฝนด้วย Seq to Seq โดยกำหนดจำนวนรอบการฝึกฝนเป็น 75000 และจำนวนเลเยอร์ RNN เป็น 1 โดยมีขนาดซ่อนเร้น (hidden size) เท่ากับ 512

lang1 = 'eng'
lang2 = 'ind'
source, target, pairs = process_data(lang1, lang2)

randomize = random.choice(pairs)
print('random sentence {}'.format(randomize))

#print number of words
input_size = source.n_words
output_size = target.n_words
print('Input : {} Output : {}'.format(input_size, output_size))

embed_size = 256
hidden_size = 512
num_layers = 1
num_iteration = 100000

#create encoder-decoder model
encoder = Encoder(input_size, hidden_size, embed_size, num_layers)
decoder = Decoder(output_size, hidden_size, embed_size, num_layers)

model = Seq2Seq(encoder, decoder, device).to(device)

#print model
print(encoder)
print(decoder)

model = trainModel(model, source, target, pairs, num_iteration)
evaluateRandomly(model, source, target, pairs)

อย่างที่คุณเห็น ประโยคที่คาดเดาของเราไม่ตรงกันมากนัก ดังนั้นเพื่อให้ได้รับความแม่นยำมากขึ้น คุณจะต้องฝึกฝนโดยใช้ข้อมูลมากขึ้นและพยายามเพิ่มการวนซ้ำและจำนวนเลเยอร์โดยใช้ Sequence เพื่อการเรียนรู้ตามลำดับ

random sentence ['tom is finishing his work', 'tom sedang menyelesaikan pekerjaannya']
Input : 3551 Output : 4253
Encoder(
  (embedding): Embedding(3551, 256)
  (gru): GRU(256, 512)
)
Decoder(
  (embedding): Embedding(4253, 256)
  (gru): GRU(256, 512)
  (out): Linear(in_features=512, out_features=4253, bias=True)
  (softmax): LogSoftmax()
)
5000 4.0906
10000 3.9129
15000 3.8171
20000 3.8369
25000 3.8199
30000 3.7957
75000 3.7044

คำถามที่พบบ่อย

โมเดล Seq2Seq แปลงลำดับหนึ่งไปเป็นอีกลำดับหนึ่ง จึงเหมาะสำหรับงานที่ความยาวของข้อมูลเข้าและข้อมูลออกแตกต่างกัน การใช้งานทั่วไป ได้แก่ การแปลภาษาด้วยเครื่อง การสรุปข้อความ การตอบคำถาม การรู้จำเสียงพูด และการสร้างคำตอบสำหรับแชทบอท

ตัวเข้ารหัสจะอ่านลำดับอินพุตและบีบอัดให้เป็นเวกเตอร์สถานะที่ซ่อนอยู่ ตัวถอดรหัสจะนำสถานะนั้นมาสร้างลำดับเอาต์พุตทีละโทเค็น โดยนำการคาดการณ์แต่ละครั้งมาใช้ซ้ำเป็นอินพุตถัดไป

GRU หรือ Gated Recurrent Unit จัดการกับลำดับข้อมูลที่ยาวได้ดีกว่า RNN ทั่วไป โดยใช้เกตแบบรีเซ็ตและอัปเดตเพื่อควบคุมหน่วยความจำ นอกจากนี้ยังมีน้ำหนักเบากว่า LSTM ทำให้การฝึกฝนทำได้เร็วขึ้นบนฮาร์ดแวร์ที่มีสเปคไม่สูงมากนัก

ในระหว่างการฝึก โมเดลจะป้อนคำเป้าหมายที่แท้จริงเข้าไปในตัวถอดรหัส แทนที่จะเป็นการคาดการณ์ของโมเดลเอง โดยควบคุมด้วย teacher_forcing_ratio ซึ่งจะช่วยเร่งการบรรลุผลลัพธ์ที่ต้องการและลดการสะสมข้อผิดพลาดตลอดลำดับเอาต์พุต

โทเค็น SOS (start of sequence) บอกให้ตัวถอดรหัสเริ่มสร้าง และ EOS (จุดสิ้นสุดของลำดับ) โทเค็นจะทำเครื่องหมายจุดสิ้นสุดของประโยค ทั้งหมดนี้ช่วยให้โมเดลสามารถจัดการกับอินพุตและเอาต์พุตที่มีความยาวแปรผันได้

ในแมชชีนเลิร์นนิง seq2seq เป็นวิธีการหลักแบบกำกับดูแลสำหรับการแปลงลำดับดีเอ็นเอ สร้างขึ้นที่นี่ด้วย PyTorchโดยระบบจะเรียนรู้การจับคู่ประโยคต้นฉบับกับประโยคเป้าหมาย และขยายไปสู่ระบบสรุปใจความและบทสนทนา

โมเดล AI สมัยใหม่ เช่น GPT ใช้ Transformer ซึ่งพัฒนามาจากแนวคิดตัวเข้ารหัส-ตัวถอดรหัส seq2seq บวกกับกลไก Attention ปัจจุบัน Transformer เป็นผู้นำในงานส่วนใหญ่ แต่การเรียนรู้ seq2seq แบบคลาสสิกยังคงอธิบายถึงพื้นฐานที่ระบบเหล่านี้สร้างขึ้นได้

ความแม่นยำต่ำมักเกิดจากข้อมูลฝึกฝนน้อยเกินไปหรือจำนวนรอบการฝึกฝนน้อยเกินไป ควรเพิ่มขนาดชุดข้อมูล เพิ่มจำนวนรอบการฝึกฝนและเลเยอร์ RNN และพิจารณาเพิ่มกลไกความสนใจ (attention mechanism) เพื่อปรับปรุงคุณภาพการแปลในประโยคที่ยาวขึ้น

สรุปโพสต์นี้ด้วย: