Generation Config

Generation Config

gonewilds


Temperature: The temperature controls the degree of randomness in token selection. The temperature is used for sampling during response generation, which occurs when topP and topK are applied. Lower temperatures are good for prompts that require a more deterministic or less open-ended response, while higher temperatures can lead to more diverse or creative results. A temperature of 0 is deterministic, meaning that the highest probability response is always selected.


topP: The topP parameter changes how the model selects tokens for output. Tokens are selected from the most to least probable until the sum of their probabilities equals the topP value. For example, if tokens A, B, and C have a probability of 0.3, 0.2, and 0.1 and the topP value is 0.5, then the model will select either A or B as the next token by using the temperature and exclude C as a candidate. The default topP value is 0.95.


topK: The topK parameter changes how the model selects tokens for output. A topK of 1 means the selected token is the most probable among all the tokens in the model's vocabulary (also called greedy decoding), while a topK of 3 means that the next token is selected from among the 3 most probable using the temperature. For each token selection step, the topK tokens with the highest probabilities are sampled. Tokens are then further filtered based on topP with the final token selected using temperature sampling.


Max output tokens: Specifies the maximum number of tokens that can be generated in the response. A token is approximately four characters. 100 tokens correspond to roughly 60-80 words.



For reducing repetitive responses, you'll want to use positive values for both presencePenalty and frequencyPenalty. These will discourage the model from reusing the same tokens frequently. Here are some general starting points:


Presence Penalty: Set this to a small positive value, like 0.5. This will lightly discourage repeated use of specific tokens.

Frequency Penalty: Set this slightly higher, around 0.7 to 1.0. This will add a stronger penalty each time a token is reused, helping to encourage a wider vocabulary in responses.



Report Page