Documentation
¶
Overview ¶
Package hfgo provides Go bindings for the Hugging Face Inference API.
Design notes:
- Clients are immutable; options are fixed at creation time, and each call snapshots them.
- Clients are safe for concurrent use by default or when configured with immutable or synchronized dependencies.
- Per-request options can override client defaults for a single call.
- Request options are applied by value with defensive header copies; contexts and HTTP clients are shared.
- HTTP client injection uses a value factory; return a fresh client value to avoid shared state.
- The SDK favors upstream feature parity and uses DTOs closely aligned to the API; breaking changes are possible as the upstream API evolves.
- WithDefaultHTTPClient restores the default client; a nil factory is treated as a configuration error.
- Client.Raw() returns the RawService escape hatch for arbitrary endpoints, exposing both error-interpreting and raw request paths (Do vs DoRaw).
- DTO validation is enforced during JSON marshal/unmarshal. Invalid request payloads surface as configuration errors. For responses, invalid content type surfaces as validation errors, while malformed JSON surfaces as serialization errors.
- Concurrency assumes externally supplied objects (for example, transports) are not mutated after use unless callers provide their own synchronization.
- Request DTOs are passed to Client methods by value and the SDK never mutates the caller's payload.
- A request and the nested data it references must be treated as read-only while a call is in flight. For concurrent invocation, pass a defensive copy per call (for example, go client.Chat(req.Clone(), ...)) or build a fresh request per call. Every request DTO provides a deep Clone method.
Index ¶
- Constants
- func UserAgent() string
- type APIError
- type ChatChoice
- type ChatCompletionMessage
- type ChatFunctionCall
- type ChatFunctionDefinition
- type ChatFunctionName
- type ChatImageURL
- type ChatJSONSchemaConfig
- type ChatLogProb
- type ChatLogProbs
- type ChatMessage
- type ChatMessageChunk
- type ChatMessageContent
- type ChatRequest
- type ChatResponse
- type ChatResponseFormat
- type ChatStream
- type ChatStreamChoice
- type ChatStreamDelta
- type ChatStreamFunction
- type ChatStreamOptions
- type ChatStreamResponse
- type ChatStreamToolCall
- type ChatTool
- type ChatToolCall
- type ChatToolCallOutput
- type ChatToolChoice
- type ChatTopLogProb
- type ChatUsage
- type Client
- func (c Client) AnswerQuestion(req QuestionAnsweringRequest, opts ...Option) ([]QuestionAnswering, error)
- func (c Client) AnswerTableQuestion(req TableQuestionAnsweringRequest, opts ...Option) (TableQuestionAnswer, error)
- func (c Client) Chat(req ChatRequest, opts ...Option) (ChatResponse, error)
- func (c Client) ChatStream(req ChatRequest, opts ...Option) (*ChatStream, error)
- func (c Client) ClassifyText(req TextClassificationRequest, opts ...Option) ([]TextClassification, error)
- func (c Client) ClassifyTextBatch(req TextClassificationBatchRequest, opts ...Option) ([][]TextClassification, error)
- func (c Client) ClassifyTokens(req TokenClassificationRequest, opts ...Option) ([]TokenClassification, error)
- func (c Client) ClassifyTokensBatch(req TokenClassificationBatchRequest, opts ...Option) ([][]TokenClassification, error)
- func (c Client) FillMask(req FillMaskRequest, opts ...Option) ([]FillMaskPrediction, error)
- func (c Client) FillMaskBatch(req FillMaskBatchRequest, opts ...Option) ([][]FillMaskPrediction, error)
- func (c Client) Raw() RawService
- func (c Client) Summarize(req SummarizationRequest, opts ...Option) ([]Summarization, error)
- func (c Client) SummarizeBatch(req SummarizationBatchRequest, opts ...Option) ([]Summarization, error)
- func (c Client) Translate(req TranslationRequest, opts ...Option) ([]Translation, error)
- func (c Client) TranslateBatch(req TranslationBatchRequest, opts ...Option) ([]Translation, error)
- func (c Client) ZeroShotClassifyText(req ZeroShotTextClassificationRequest, opts ...Option) ([]ZeroShotTextClassification, error)
- func (c Client) ZeroShotClassifyTextBatch(req ZeroShotTextClassificationBatchRequest, opts ...Option) ([][]ZeroShotTextClassification, error)
- type FillMaskBatchRequest
- type FillMaskParameters
- type FillMaskPrediction
- type FillMaskRequest
- type MessageChunkType
- type Option
- func WithBaseURL(u string) Option
- func WithContext(ctx context.Context) Option
- func WithDefaultHTTPClient() Option
- func WithDefaultHeader(key, value string) Option
- func WithHTTPClientFactory(factory func() http.Client) Option
- func WithHeader(key, value string) Option
- func WithHeaders(h http.Header) Option
- func WithMaxResponseBodyBytes(n int64) Option
- func WithModel(m string) Option
- func WithProvider(p string) Option
- func WithToken(t string) Option
- func WithUserAgentSuffix(s string) Option
- type QuestionAnswering
- type QuestionAnsweringInput
- type QuestionAnsweringParameters
- type QuestionAnsweringRequest
- type RawEvent
- type RawService
- func (r RawService) Do(requestBody []byte, method string, path string, opts ...Option) (*http.Response, error)
- func (r RawService) DoRaw(requestBody []byte, method string, path string, opts ...Option) (*http.Response, error)
- func (r RawService) DoRawReader(requestBody io.Reader, method string, path string, opts ...Option) (*http.Response, error)
- func (r RawService) DoReader(requestBody io.Reader, method string, path string, opts ...Option) (*http.Response, error)
- func (r RawService) Stream(requestBody []byte, method string, path string, opts ...Option) (*RawStream, error)
- func (r RawService) StreamRaw(requestBody []byte, method string, path string, opts ...Option) (*RawStream, error)
- func (r RawService) StreamRawReader(requestBody io.Reader, method string, path string, opts ...Option) (*RawStream, error)
- func (r RawService) StreamReader(requestBody io.Reader, method string, path string, opts ...Option) (*RawStream, error)
- type RawStream
- type ResponseFormatType
- type SDKError
- type SDKErrorKind
- type Summarization
- type SummarizationBatchRequest
- type SummarizationParameters
- type SummarizationRequest
- type TableQuestionAnswer
- type TableQuestionAnsweringInput
- type TableQuestionAnsweringParameters
- type TableQuestionAnsweringRequest
- type TextClassification
- type TextClassificationBatchRequest
- type TextClassificationParameters
- type TextClassificationRequest
- type TokenClassification
- type TokenClassificationBatchRequest
- type TokenClassificationParameters
- type TokenClassificationRequest
- type ToolChoiceMode
- type Translation
- type TranslationBatchRequest
- type TranslationParameters
- type TranslationRequest
- type ZeroShotTextClassification
- type ZeroShotTextClassificationBatchRequest
- type ZeroShotTextClassificationParameters
- type ZeroShotTextClassificationRequest
Constants ¶
const ( // SDKErrorKindValidation indicates a validation error in API responses. SDKErrorKindValidation = hferrors.SDKErrorKindValidation // SDKErrorKindConfiguration indicates invalid or missing configuration. SDKErrorKindConfiguration = hferrors.SDKErrorKindConfiguration // SDKErrorKindSerialization indicates a serialization or deserialization error. SDKErrorKindSerialization = hferrors.SDKErrorKindSerialization // SDKErrorKindTransport indicates a transport-layer failure. SDKErrorKindTransport = hferrors.SDKErrorKindTransport // SDKErrorKindInternal indicates an internal SDK error. SDKErrorKindInternal = hferrors.SDKErrorKindInternal )
const ( // SummarizationTruncationDoNotTruncate keeps the input as-is without any truncation. SummarizationTruncationDoNotTruncate = "do_not_truncate" // SummarizationTruncationLongestFirst truncates the longest side first when the input exceeds the model's maximum length. SummarizationTruncationLongestFirst = "longest_first" // SummarizationTruncationOnlyFirst trims the input on the first (left) side when it exceeds the model's maximum length. SummarizationTruncationOnlyFirst = "only_first" // SummarizationTruncationOnlySecond trims the input on the second (right) side when it exceeds the model's maximum length. SummarizationTruncationOnlySecond = "only_second" )
const ( // TableQuestionAnsweringPaddingDoNotPad does not pad the input. TableQuestionAnsweringPaddingDoNotPad = "do_not_pad" // TableQuestionAnsweringPaddingLongest pads to the longest sequence in the batch. TableQuestionAnsweringPaddingLongest = "longest" // TableQuestionAnsweringPaddingMaxLength pads to the maximum length. TableQuestionAnsweringPaddingMaxLength = "max_length" )
const ( // TextClassificationFuncSigmoid applies a sigmoid to each score independently. // Useful for multi-label classification tasks, where multiple classes may apply simultaneously. TextClassificationFuncSigmoid = "sigmoid" // TextClassificationFuncSoftmax normalizes scores into a probability distribution summing to 1. // Useful for single-label multi-class classification tasks, where exactly one class applies. TextClassificationFuncSoftmax = "softmax" // TextClassificationFuncNone returns raw scores without any transformation applied. TextClassificationFuncNone = "none" )
const ( // TokenClassificationAggregationNone does not aggregate tokens. // Each token is classified individually. TokenClassificationAggregationNone = "none" // TokenClassificationAggregationSimple groups consecutive tokens with the same // label into a single entity. TokenClassificationAggregationSimple = "simple" // TokenClassificationAggregationFirst is similar to "simple", but also preserves // word integrity using the label predicted for the first token in a word. TokenClassificationAggregationFirst = "first" // TokenClassificationAggregationAverage is similar to "simple", but also preserves // word integrity using the label with the highest score, averaged across the // word's tokens. TokenClassificationAggregationAverage = "average" // TokenClassificationAggregationMax is similar to "simple", but also preserves // word integrity using the label with the highest score across the word's tokens. TokenClassificationAggregationMax = "max" )
const ( // TranslationTruncationDoNotTruncate keeps the input as-is without any truncation. TranslationTruncationDoNotTruncate = "do_not_truncate" // TranslationTruncationLongestFirst truncates the longest side first when the input exceeds the model's maximum length. TranslationTruncationLongestFirst = "longest_first" // TranslationTruncationOnlyFirst trims the input on the first (left) side when it exceeds the model's maximum length. TranslationTruncationOnlyFirst = "only_first" // TranslationTruncationOnlySecond trims the input on the second (right) side when it exceeds the model's maximum length. TranslationTruncationOnlySecond = "only_second" )
const EndpointChatCompletion = "/v1/chat/completions"
EndpointChatCompletion specifies the chat completion endpoint.
const Version = sdkversion.Version
Version is the current version of the hfgo SDK. This follows semantic versioning (semver.org).
Variables ¶
This section is empty.
Functions ¶
Types ¶
type APIError ¶
APIError represents an error returned by the HuggingFace API. It includes the HTTP status code, error message, response body, and request ID if available.
Users can type-assert errors to *APIError to access additional error information and helper methods:
if apiErr, ok := err.(*hfgo.APIError); ok {
if apiErr.IsAuthenticationError() {
// Handle authentication error
}
}
type ChatChoice ¶
type ChatChoice struct {
// Required.
FinishReason string `json:"finish_reason"`
// Required.
Index int `json:"index"`
LogProbs *ChatLogProbs `json:"logprobs,omitempty"`
// Required.
Message ChatCompletionMessage `json:"message"`
}
ChatChoice is a single non-streaming completion choice.
type ChatCompletionMessage ¶
type ChatCompletionMessage struct {
Role string `json:"role"`
// Content is present for text responses.
Content *string `json:"content,omitempty"`
// ToolCallID is set when returning tool-specific content.
ToolCallID *string `json:"tool_call_id,omitempty"`
// ToolCalls is present for tool call responses.
ToolCalls []ChatToolCallOutput `json:"tool_calls,omitempty"`
}
ChatCompletionMessage is a message returned by the model. It is either a text message (Content) or a tool call message (ToolCalls).
func (*ChatCompletionMessage) UnmarshalJSON ¶
func (m *ChatCompletionMessage) UnmarshalJSON(data []byte) error
UnmarshalJSON enforces the union shape for ChatCompletionMessage.
type ChatFunctionCall ¶
type ChatFunctionCall struct {
// Required.
Name string `json:"name"`
// Required.
Arguments string `json:"arguments"`
Description *string `json:"description,omitempty"`
}
ChatFunctionCall represents a tool function call with arguments.
func (*ChatFunctionCall) Clone ¶
func (f *ChatFunctionCall) Clone() ChatFunctionCall
Clone returns a deep defensive copy of the function call.
func (ChatFunctionCall) MarshalJSON ¶
func (f ChatFunctionCall) MarshalJSON() ([]byte, error)
MarshalJSON enforces required fields on ChatFunctionCall.
func (*ChatFunctionCall) UnmarshalJSON ¶
func (f *ChatFunctionCall) UnmarshalJSON(data []byte) error
UnmarshalJSON enforces required fields on ChatFunctionCall.
type ChatFunctionDefinition ¶
type ChatFunctionDefinition struct {
// Required.
Name string `json:"name"`
// A description of what the function does.
Description *string `json:"description,omitempty"`
// JSON schema describing function parameters.
Parameters json.RawMessage `json:"parameters,omitempty"`
}
ChatFunctionDefinition describes a callable function.
func (*ChatFunctionDefinition) Clone ¶
func (f *ChatFunctionDefinition) Clone() ChatFunctionDefinition
Clone returns a deep defensive copy of the function definition.
func (ChatFunctionDefinition) MarshalJSON ¶
func (f ChatFunctionDefinition) MarshalJSON() ([]byte, error)
MarshalJSON enforces required fields on ChatFunctionDefinition.
type ChatFunctionName ¶
type ChatFunctionName struct {
// Required.
Name string `json:"name"`
}
ChatFunctionName identifies a tool function by name.
func (*ChatFunctionName) Clone ¶
func (f *ChatFunctionName) Clone() ChatFunctionName
Clone returns a defensive copy of the function name.
func (ChatFunctionName) MarshalJSON ¶
func (f ChatFunctionName) MarshalJSON() ([]byte, error)
MarshalJSON enforces required fields on ChatFunctionName.
type ChatImageURL ¶
type ChatImageURL struct {
// Required.
URL string `json:"url"`
}
ChatImageURL contains the URL for an image chunk.
func (*ChatImageURL) Clone ¶
func (u *ChatImageURL) Clone() ChatImageURL
Clone returns a defensive copy of the image URL.
func (ChatImageURL) MarshalJSON ¶
func (u ChatImageURL) MarshalJSON() ([]byte, error)
MarshalJSON enforces the required URL on ChatImageURL.
type ChatJSONSchemaConfig ¶
type ChatJSONSchemaConfig struct {
// The name of the response format.
// Required.
Name string `json:"name"`
// A description of what the response format is for.
Description *string `json:"description,omitempty"`
// The schema for the response format as a JSON Schema object.
// Learn how to build JSON schemas at https://json-schema.org/.
Schema json.RawMessage `json:"schema,omitempty"`
// Whether to enable strict schema adherence.
Strict *bool `json:"strict,omitempty"`
}
ChatJSONSchemaConfig defines JSON schema response formatting.
func (*ChatJSONSchemaConfig) Clone ¶
func (c *ChatJSONSchemaConfig) Clone() ChatJSONSchemaConfig
Clone returns a deep defensive copy of the JSON schema configuration.
type ChatLogProb ¶
type ChatLogProb struct {
// Required.
Token string `json:"token"`
// Required.
LogProb float64 `json:"logprob"`
// Required.
TopLogProbs []ChatTopLogProb `json:"top_logprobs"`
}
ChatLogProb contains logprob information for a token.
type ChatLogProbs ¶
type ChatLogProbs struct {
// Required.
Content []ChatLogProb `json:"content"`
}
ChatLogProbs contains per-token log probabilities.
type ChatMessage ¶
type ChatMessage struct {
// Role of the message author (for example: system, user, assistant, tool).
// Required.
Role string `json:"role"`
// Optional name for the participant.
Name *string `json:"name,omitempty"`
// Content may be a string or a list of content chunks.
// Either Content or ToolCalls should be supplied.
Content ChatMessageContent `json:"content"`
// ToolCalls is used instead of Content when providing tool call messages.
ToolCalls []ChatToolCall `json:"tool_calls,omitempty"`
}
ChatMessage represents a single chat message.
func (*ChatMessage) Clone ¶
func (m *ChatMessage) Clone() ChatMessage
Clone returns a deep defensive copy of the message.
func (ChatMessage) MarshalJSON ¶
func (m ChatMessage) MarshalJSON() ([]byte, error)
MarshalJSON enforces the union shape for ChatMessage.
type ChatMessageChunk ¶
type ChatMessageChunk struct {
// Required when Type is "text".
Text *string `json:"text,omitempty"`
// Required when Type is "image_url".
ImageURL *ChatImageURL `json:"image_url,omitempty"`
// Required. Possible values: text, image_url.
Type MessageChunkType `json:"type"`
}
ChatMessageChunk represents a content chunk (text or image URL). Type must be "text" with Text set, or "image_url" with ImageURL set.
func (*ChatMessageChunk) Clone ¶
func (c *ChatMessageChunk) Clone() ChatMessageChunk
Clone returns a deep defensive copy of the chunk.
func (ChatMessageChunk) MarshalJSON ¶
func (c ChatMessageChunk) MarshalJSON() ([]byte, error)
MarshalJSON enforces the union shape for ChatMessageChunk.
type ChatMessageContent ¶
type ChatMessageContent struct {
// Text holds plain string content.
Text *string `json:"-"`
// Chunks holds structured content chunks.
Chunks []ChatMessageChunk `json:"-"`
}
ChatMessageContent can be a string or []ChatMessageChunk. Use a string for pure text content or a chunk list for multimodal content.
func (*ChatMessageContent) Clone ¶
func (c *ChatMessageContent) Clone() ChatMessageContent
Clone returns a deep defensive copy of the content.
func (ChatMessageContent) MarshalJSON ¶
func (c ChatMessageContent) MarshalJSON() ([]byte, error)
MarshalJSON enforces the union shape for ChatMessageContent.
type ChatRequest ¶
type ChatRequest struct {
// Model to use for the chat completion.
// Required.
Model *string `json:"model,omitempty"`
// Number between -2.0 and 2.0. Positive values penalize new tokens based
// on their existing frequency in the text so far, decreasing the model's
// likelihood to repeat the same line verbatim.
FrequencyPenalty *float64 `json:"frequency_penalty,omitempty"`
// Whether to return log probabilities of the output tokens or not.
// If true, returns the log probabilities of each output token returned
// in the content of message.
LogProbs *bool `json:"logprobs,omitempty"`
// The maximum number of tokens that can be generated in the chat completion.
// Default: 1024. Minimum: 0.
MaxTokens *int `json:"max_tokens,omitempty"`
// A list of messages comprising the conversation so far.
// Required.
Messages []ChatMessage `json:"messages"`
// Number between -2.0 and 2.0. Positive values penalize new tokens based
// on whether they appear in the text so far, increasing the model's
// likelihood to talk about new topics.
PresencePenalty *float64 `json:"presence_penalty,omitempty"`
// Response format configuration. Known types: text, json_schema, json_object.
// Non-empty provider-specific values are also accepted.
ResponseFormat *ChatResponseFormat `json:"response_format,omitempty"`
// Seed for deterministic sampling.
// Minimum: 0.
Seed *int64 `json:"seed,omitempty"`
// Up to 4 sequences where the API will stop generating further tokens.
Stop []string `json:"stop,omitempty"`
// If true, generated tokens are returned as a stream using SSE.
// For more information about streaming, see:
// https://huggingface.co/docs/text-generation-inference/conceptual/streaming
Stream *bool `json:"stream,omitempty"`
// Stream options for SSE responses.
StreamOptions *ChatStreamOptions `json:"stream_options,omitempty"`
// Sampling temperature to use, between 0 and 2.
// We generally recommend altering this or TopP but not both.
Temperature *float64 `json:"temperature,omitempty"`
// Tool choice behavior. Known values: auto, none, required, or
// {"function":{"name":"..."}}. Meanings:
// - auto: model can pick between generating a message or calling tools.
// - none: model will not call any tool and instead generates a message.
// - required: model must call one or more tools.
// Non-empty provider-specific values are also accepted.
ToolChoice *ChatToolChoice `json:"tool_choice,omitempty"`
// A prompt to be appended before the tools.
ToolPrompt *string `json:"tool_prompt,omitempty"`
// A list of tools the model may call.
// Currently, only functions are supported as a tool. Use this to provide
// a list of functions the model may generate JSON inputs for.
Tools []ChatTool `json:"tools,omitempty"`
// Number of most likely tokens to return at each token position.
// LogProbs must be true if this is set.
// An integer between 0 and 5.
TopLogProbs *int `json:"top_logprobs,omitempty"`
// Nucleus sampling probability mass.
// For example, 0.1 means only the tokens comprising the top 10% probability
// mass are considered.
TopP *float64 `json:"top_p,omitempty"`
}
ChatRequest represents a completion request for the chat API. Output type depends on the Stream parameter.
func (*ChatRequest) Clone ¶
func (r *ChatRequest) Clone() ChatRequest
Clone returns a deep defensive copy of the request. All reference fields — slices, maps, pointers, and JSON blobs — are copied so the returned value shares no backing storage with the receiver. This makes it safe to reuse a template request across concurrent calls, e.g. go client.Chat(req.Clone(), ...).
func (ChatRequest) MarshalJSON ¶
func (r ChatRequest) MarshalJSON() ([]byte, error)
MarshalJSON enforces required fields for ChatRequest.
type ChatResponse ¶
type ChatResponse struct {
// Required.
ID string `json:"id"`
// Unix timestamp in seconds.
// Required.
Created int64 `json:"created"`
// Required.
Model string `json:"model"`
// Required.
SystemFingerprint string `json:"system_fingerprint"`
// Required.
Choices []ChatChoice `json:"choices"`
// Required.
Usage ChatUsage `json:"usage"`
}
ChatResponse represents a non-streaming completion response from the chat API. This is returned when Stream is false (the default).
type ChatResponseFormat ¶
type ChatResponseFormat struct {
// Known type values: text, json_schema, json_object.
// Non-empty provider-specific values are also accepted.
// For json_schema, JSONSchema is required.
Type ResponseFormatType `json:"type"`
JSONSchema *ChatJSONSchemaConfig `json:"json_schema,omitempty"`
}
ChatResponseFormat configures the response format.
func (*ChatResponseFormat) Clone ¶
func (r *ChatResponseFormat) Clone() ChatResponseFormat
Clone returns a deep defensive copy of the response format.
func (ChatResponseFormat) MarshalJSON ¶
func (r ChatResponseFormat) MarshalJSON() ([]byte, error)
MarshalJSON enforces the union shape for ChatResponseFormat.
type ChatStream ¶
type ChatStream struct {
// contains filtered or unexported fields
}
ChatStream wraps a streaming chat completion response.
func (*ChatStream) Close ¶
func (c *ChatStream) Close() error
Close releases the underlying stream resources.
func (*ChatStream) Recv ¶
func (c *ChatStream) Recv(ctx context.Context) (ChatStreamResponse, error)
Recv blocks until the next streaming chunk arrives or the context is done.
type ChatStreamChoice ¶
type ChatStreamChoice struct {
// Required.
Delta ChatStreamDelta `json:"delta"`
FinishReason *string `json:"finish_reason,omitempty"`
// Required.
Index int `json:"index"`
LogProbs *ChatLogProbs `json:"logprobs,omitempty"`
}
ChatStreamChoice is a single streaming completion choice.
type ChatStreamDelta ¶
type ChatStreamDelta struct {
// Content is present for text deltas.
Content *string `json:"content,omitempty"`
// Role may be included with the first delta.
Role *string `json:"role,omitempty"`
// ToolCallID may be included for tool-specific content.
ToolCallID *string `json:"tool_call_id,omitempty"`
// ToolCalls is present for tool call deltas.
ToolCalls []ChatStreamToolCall `json:"tool_calls,omitempty"`
}
ChatStreamDelta holds incremental updates for a stream. Deltas may include content/role/tool_call_id or role/tool_calls.
func (*ChatStreamDelta) UnmarshalJSON ¶
func (d *ChatStreamDelta) UnmarshalJSON(data []byte) error
UnmarshalJSON enforces the union shape for ChatStreamDelta.
type ChatStreamFunction ¶
ChatStreamFunction represents a streamed function call.
type ChatStreamOptions ¶
type ChatStreamOptions struct {
// If set, an additional chunk is streamed before the [DONE] message
// showing overall usage. The usage field on this chunk shows the token
// usage statistics for the entire request, and the choices field will
// always be an empty array. All other chunks include a usage field with
// a null value.
IncludeUsage *bool `json:"include_usage,omitempty"`
}
ChatStreamOptions configures streaming behavior.
func (*ChatStreamOptions) Clone ¶
func (o *ChatStreamOptions) Clone() ChatStreamOptions
Clone returns a deep defensive copy of the stream options.
type ChatStreamResponse ¶
type ChatStreamResponse struct {
ID string `json:"id"`
// Unix timestamp in seconds.
Created int64 `json:"created"`
Model string `json:"model"`
SystemFingerprint string `json:"system_fingerprint"`
Choices []ChatStreamChoice `json:"choices"`
Usage *ChatUsage `json:"usage,omitempty"`
}
ChatStreamResponse represents a streaming response chunk. This is returned when Stream is true.
type ChatStreamToolCall ¶
type ChatStreamToolCall struct {
ID string `json:"id,omitempty"`
Type string `json:"type,omitempty"`
Index int `json:"index"`
Function ChatStreamFunction `json:"function"`
}
ChatStreamToolCall represents a tool call within a streaming delta.
type ChatTool ¶
type ChatTool struct {
// Tool type (currently only function is supported).
Type string `json:"type"`
Function ChatFunctionDefinition `json:"function"`
}
ChatTool represents a tool definition provided to the model.
func (ChatTool) MarshalJSON ¶
MarshalJSON enforces the tool shape for ChatTool.
type ChatToolCall ¶
type ChatToolCall struct {
// Tool call ID.
// Required.
ID string `json:"id"`
// Tool call type (for example: function).
// Required.
Type string `json:"type"`
// Required.
Function ChatFunctionCall `json:"function"`
}
ChatToolCall represents a tool call in a message.
func (*ChatToolCall) Clone ¶
func (c *ChatToolCall) Clone() ChatToolCall
Clone returns a deep defensive copy of the tool call.
func (ChatToolCall) MarshalJSON ¶
func (c ChatToolCall) MarshalJSON() ([]byte, error)
MarshalJSON enforces the tool call shape for ChatToolCall.
type ChatToolCallOutput ¶
type ChatToolCallOutput struct {
// Required.
ID string `json:"id"`
// Required.
Type string `json:"type"`
// Required.
Function ChatFunctionCall `json:"function"`
}
ChatToolCallOutput represents a tool call in a response message.
func (*ChatToolCallOutput) UnmarshalJSON ¶
func (c *ChatToolCallOutput) UnmarshalJSON(data []byte) error
UnmarshalJSON enforces the tool call shape for ChatToolCallOutput.
type ChatToolChoice ¶
type ChatToolChoice struct {
// Mode is a tool choice mode. Known values: auto, none, required.
// Non-empty provider-specific values are also accepted.
Mode *ToolChoiceMode `json:"-"`
// Function selects a specific tool function by name.
Function *ChatFunctionName `json:"-"`
}
ChatToolChoice represents the tool choice union type. Either Mode is set to auto/none/required, or Function is set.
func (*ChatToolChoice) Clone ¶
func (t *ChatToolChoice) Clone() ChatToolChoice
Clone returns a deep defensive copy of the tool choice.
func (ChatToolChoice) MarshalJSON ¶
func (t ChatToolChoice) MarshalJSON() ([]byte, error)
MarshalJSON enforces the union shape for ChatToolChoice.
type ChatTopLogProb ¶
type ChatTopLogProb struct {
// Required.
Token string `json:"token"`
// Required.
LogProb float64 `json:"logprob"`
}
ChatTopLogProb is a top log probability entry for a token position.
type ChatUsage ¶
type ChatUsage struct {
// Required.
CompletionTokens int `json:"completion_tokens"`
// Required.
PromptTokens int `json:"prompt_tokens"`
// Required.
TotalTokens int `json:"total_tokens"`
}
ChatUsage contains token usage statistics.
type Client ¶
type Client struct {
// contains filtered or unexported fields
}
Client represents a HuggingFace API client with configured request options. Client instances are immutable; options are fixed at creation time and never mutated. This keeps client usage safe across goroutines and avoids surprises from mutable state. If options include externally-owned pointers, callers must avoid mutating them after creation or ensure their own synchronization. RawService captures a snapshot of these options when created.
func NewClient ¶
NewClient creates a new Client instance with the provided request options. If no options are provided, default options will be used. Clients are immutable; to change options, create a new Client to keep calls deterministic.
func (Client) AnswerQuestion ¶
func (c Client) AnswerQuestion( req QuestionAnsweringRequest, opts ...Option, ) ([]QuestionAnswering, error)
AnswerQuestion sends a question answering request and returns the answers.
The request must include both a question and a context. The model will identify the answer to the question within the provided context.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) AnswerTableQuestion ¶
func (c Client) AnswerTableQuestion( req TableQuestionAnsweringRequest, opts ...Option, ) (TableQuestionAnswer, error)
AnswerTableQuestion sends a table question answering request and returns the answer.
The request must include both a question and a table. The model will identify the answer to the question within the provided table data.
NOTE: The HuggingFace API returns a bare JSON object for table question answering, not an array — despite the upstream schema declaring an array response. This method returns a single TableQuestionAnswer to match the actual API behavior.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) Chat ¶
func (c Client) Chat(req ChatRequest, opts ...Option) (ChatResponse, error)
Chat sends a chat completion request and returns a chat completion response.
The request is passed by value and the SDK never mutates the received payload. The value copy shares the request's nested data (slices, maps, and pointed-to values) with the caller, so the caller must treat the request and the data it references as read-only while a call is in flight.
Concurrency:
- A single Client is safe for concurrent use.
- Reusing one request across sequential, fully-awaited calls is safe.
- To invoke the same template request from multiple goroutines, pass a defensive copy per call, e.g. go client.Chat(req.Clone(), ...), or build a fresh request per call.
Model Precedence: The Model field is resolved with the following precedence (highest to lowest):
- ChatRequest.Model field (if non-nil and non-empty)
- Per-request options Model override
- Client-level Model option
Provider Precedence: The Provider field is applied as a fallback only if the resolved Model does not already contain a provider (indicated by ":" in the model string). If the Model is in the format "model:provider", the Provider option is ignored.
For example:
- Model="mistral-7b", Provider="huggingface" → "mistral-7b:huggingface"
- Model="mistral-7b:huggingface", Provider="mistral" → "mistral-7b:huggingface" (Provider ignored)
- Model="mistral-7b:huggingface", Provider="" → "mistral-7b:huggingface"
Behavior:
- Returns a configuration error if the request is missing a model or messages.
- Returns a configuration error if *req.Stream is true; use ChatStream for streaming.
func (Client) ChatStream ¶
func (c Client) ChatStream(req ChatRequest, opts ...Option) (*ChatStream, error)
ChatStream sends a chat completion request and returns a streaming response. Callers should Close the returned ChatStream when finished so the underlying HTTP connection and decoder goroutine are released promptly.
The request is passed by value and the SDK never mutates the received payload. The value copy shares the request's nested data (slices, maps, and pointed-to values) with the caller, so the caller must treat the request and the data it references as read-only while a call is in flight.
Concurrency:
- A single Client is safe for concurrent use.
- Reusing one request across sequential, fully-awaited calls is safe.
- To invoke the same template request from multiple goroutines, pass a defensive copy per call, e.g. go client.ChatStream(req.Clone(), ...), or build a fresh request per call.
Model Precedence: The Model field is resolved with the following precedence (highest to lowest):
- ChatRequest.Model field (if non-nil and non-empty)
- Per-request options Model override
- Client-level Model option
Provider Precedence: The Provider field is applied as a fallback only if the resolved Model does not already contain a provider (indicated by ":" in the model string). If the Model is in the format "model:provider", the Provider option is ignored.
For example:
- Model="mistral-7b", Provider="huggingface" → "mistral-7b:huggingface"
- Model="mistral-7b:huggingface", Provider="mistral" → "mistral-7b:huggingface" (Provider ignored)
- Model="mistral-7b:huggingface", Provider="" → "mistral-7b:huggingface"
Behavior:
- Returns a configuration error if the request is missing a model or messages.
- Always sends the request with streaming enabled.
func (Client) ClassifyText ¶
func (c Client) ClassifyText( req TextClassificationRequest, opts ...Option, ) ([]TextClassification, error)
ClassifyText sends a text classification request and returns the text classification response for a single input.
For multiple classification inputs, use ClassifyTextBatch.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) ClassifyTextBatch ¶
func (c Client) ClassifyTextBatch( req TextClassificationBatchRequest, opts ...Option, ) ([][]TextClassification, error)
ClassifyTextBatch sends a text classification request for a batch of inputs and returns a list of text classification responses for each input in the batch.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.
Callers should check the length of the response list before indexing.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) ClassifyTokens ¶
func (c Client) ClassifyTokens( req TokenClassificationRequest, opts ...Option, ) ([]TokenClassification, error)
ClassifyTokens sends a token classification request and returns the token classification response for a single input.
For multiple inputs, use ClassifyTokensBatch.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) ClassifyTokensBatch ¶
func (c Client) ClassifyTokensBatch( req TokenClassificationBatchRequest, opts ...Option, ) ([][]TokenClassification, error)
ClassifyTokensBatch sends a token classification request for a batch of inputs and returns a list of token classification responses for each input in the batch.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.
Callers should check the length of the response list before indexing.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) FillMask ¶
func (c Client) FillMask(req FillMaskRequest, opts ...Option) ([]FillMaskPrediction, error)
FillMask sends a fill mask request and returns the mask filling predictions for a single input.
For multiple inputs, use FillMaskBatch.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) FillMaskBatch ¶
func (c Client) FillMaskBatch( req FillMaskBatchRequest, opts ...Option, ) ([][]FillMaskPrediction, error)
FillMaskBatch sends a fill mask request for a batch of inputs and returns a list of mask filling predictions for each input in the batch.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.
Callers should check the length of the response list before indexing.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) Raw ¶
func (c Client) Raw() RawService
Raw returns the raw HTTP request service for this client. Unlike the other endpoints, which are exposed directly as Client methods, the raw path remains namespaced under RawService: it is the advanced escape hatch for endpoints the SDK does not otherwise cover, and its several method variants are easier to discover grouped together than splashed across the Client surface.
RawService is immutable and captures a snapshot of the client options when created; it is lightweight, so prefer calling Raw() per use rather than retaining the value.
func (Client) Summarize ¶
func (c Client) Summarize(req SummarizationRequest, opts ...Option) ([]Summarization, error)
Summarize sends a summarization request and returns the summarization output for a single input.
The API always returns a list for summarization; a single input yields a one-element list rather than a bare summary object.
For multiple inputs, use SummarizeBatch.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) SummarizeBatch ¶
func (c Client) SummarizeBatch( req SummarizationBatchRequest, opts ...Option, ) ([]Summarization, error)
SummarizeBatch sends a summarization request for a batch of inputs and returns a flat list of summarization outputs, one for each input in the batch, in the same order as the inputs.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice. The response is a flat list (one summary per input) — not a nested list — consistent with how the API returns a list even for a single input.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) Translate ¶
func (c Client) Translate(req TranslationRequest, opts ...Option) ([]Translation, error)
Translate sends a translation request and returns the translation output for a single input.
The API always returns a list for translation; a single input yields a one-element list rather than a bare translation object.
For multiple inputs, use TranslateBatch.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) TranslateBatch ¶
func (c Client) TranslateBatch( req TranslationBatchRequest, opts ...Option, ) ([]Translation, error)
TranslateBatch sends a translation request for a batch of inputs and returns a flat list of translation outputs, one for each input in the batch, in the same order as the inputs.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice. The response is a flat list (one translation per input) — not a nested list — consistent with how the API returns a list even for a single input.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) ZeroShotClassifyText ¶
func (c Client) ZeroShotClassifyText( req ZeroShotTextClassificationRequest, opts ...Option, ) ([]ZeroShotTextClassification, error)
ZeroShotClassifyText sends a zero-shot text classification request and returns the zero-shot text classification response for a single input.
For multiple inputs, use ZeroShotClassifyTextBatch.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
func (Client) ZeroShotClassifyTextBatch ¶
func (c Client) ZeroShotClassifyTextBatch( req ZeroShotTextClassificationBatchRequest, opts ...Option, ) ([][]ZeroShotTextClassification, error)
ZeroShotClassifyTextBatch sends a zero-shot text classification request for a batch of inputs and returns a list of zero-shot text classification responses for each input in the batch.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.
Callers should check the length of the response list before indexing.
The Provider option is ignored for now, as hf-inference is currently the only supported provider.
type FillMaskBatchRequest ¶
type FillMaskBatchRequest struct {
// The inputs with masked tokens.
// Required.
Inputs []string `json:"inputs"`
// Additional inference parameters for mask filling
Parameters *FillMaskParameters `json:"parameters,omitempty"`
}
FillMaskBatchRequest represents a batched fill mask request to the API for multiple masked inputs.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.
func (*FillMaskBatchRequest) Clone ¶
func (r *FillMaskBatchRequest) Clone() FillMaskBatchRequest
Clone returns a deep defensive copy of the request.
type FillMaskParameters ¶
type FillMaskParameters struct {
// When passed, the model will limit the scores to the passed targets instead of looking up
// in the whole vocabulary. If the provided targets are not in the model vocab, they will be
// tokenized and the first resulting token will be used (with a warning, and that might be
// slower).
Targets []string `json:"targets,omitempty"`
// When passed, overrides the number of predictions to return.
TopK *int `json:"top_k,omitempty"`
}
FillMaskParameters specify additional inference parameters for mask filling tasks.
func (*FillMaskParameters) Clone ¶
func (p *FillMaskParameters) Clone() FillMaskParameters
Clone returns a deep defensive copy of the parameters.
type FillMaskPrediction ¶
type FillMaskPrediction struct {
// The input filled with the mask token prediction
Sequence string `json:"sequence"`
// The probability of the token prediction
Score float64 `json:"score"`
// The predicted token id (to replace the masked one).
Token int `json:"token"`
// The predicted token (to replace the masked one).
TokenStr *string `json:"token_str"`
}
FillMaskPrediction represents a mask-filling output.
type FillMaskRequest ¶
type FillMaskRequest struct {
// The input text with masked tokens.
// Required.
Input string `json:"inputs"`
// Additional inference parameters for mask filling
Parameters *FillMaskParameters `json:"parameters,omitempty"`
}
FillMaskRequest represents a fill mask inference request to the API for a single input.
func (*FillMaskRequest) Clone ¶
func (r *FillMaskRequest) Clone() FillMaskRequest
Clone returns a deep defensive copy of the request.
type MessageChunkType ¶
type MessageChunkType string
MessageChunkType enumerates supported chat message chunk types.
const ( // MessageChunkTypeText represents a text chunk. MessageChunkTypeText MessageChunkType = "text" // MessageChunkTypeImageURL represents an image_url chunk. MessageChunkTypeImageURL MessageChunkType = "image_url" )
type Option ¶
Option represents a functional option that configures client requests.
func WithBaseURL ¶
WithBaseURL returns an Option that sets the base URL for API requests. The base URL is the root endpoint for all HuggingFace API calls and must not include query parameters or fragments.
func WithContext ¶
WithContext returns an Option that sets the context for API requests. The context can be used for cancellation, timeouts, and passing request-scoped values. If a nil context is provided, the SDK will fall back to context.Background().
func WithDefaultHTTPClient ¶
func WithDefaultHTTPClient() Option
WithDefaultHTTPClient returns an Option that sets the default HTTP client.
func WithDefaultHeader ¶
WithDefaultHeader returns an Option that sets a header only if missing or empty.
func WithHTTPClientFactory ¶
WithHTTPClientFactory returns an Option that sets a http.Client created by the factory. The factory is invoked when request options are applied, so it can be used per request or at client construction time. The factory should return a fresh client value; avoid sharing mutable internals like Transport unless synchronized. If the factory is nil, the HTTP client is set to nil.
func WithHeader ¶
WithHeader returns an Option that sets a single header applied to every request.
func WithHeaders ¶
WithHeaders returns an Option that sets custom headers applied to every request, overriding any existing values for matching keys. Per-request headers can still override these values when provided.
func WithMaxResponseBodyBytes ¶
WithMaxResponseBodyBytes returns an Option that sets the maximum number of bytes read from any response body. Values <= 0 fall back to the default.
func WithModel ¶
WithModel returns an Option that sets the model to use for API requests. The model specifies which HuggingFace model should process the request.
func WithProvider ¶
WithProvider returns an Option that sets the provider for API requests. The provider specifies which inference provider should handle the request.
func WithToken ¶
WithToken returns an Option that sets the authentication token for API requests. The token is used for Bearer authentication with the HuggingFace API.
func WithUserAgentSuffix ¶
WithUserAgentSuffix returns an Option that appends a suffix to the SDK user agent string.
type QuestionAnswering ¶
type QuestionAnswering struct {
// The answer to the question.
Answer string `json:"answer"`
// The probability associated to the answer.
Score float64 `json:"score"`
// The character position in the input where the answer begins.
Start int `json:"start"`
// The character position in the input where the answer ends.
End int `json:"end"`
}
QuestionAnswering represents a question answering output.
type QuestionAnsweringInput ¶
type QuestionAnsweringInput struct {
// The question to be answered.
// Required.
Question string `json:"question"`
// The context to be used for answering the question.
// Required.
Context string `json:"context"`
}
QuestionAnsweringInput represents the input data for a question answering request. Both Question and Context are required.
type QuestionAnsweringParameters ¶
type QuestionAnsweringParameters struct {
// The number of answers to return (will be chosen by order of likelihood).
// Note that less than top_k answers may be returned if there are not enough
// options available within the context.
TopK *int `json:"top_k,omitempty"`
// If the context is too long to fit with the question for the model, it will
// be split in several chunks with some overlap. This argument controls the
// size of that overlap.
DocStride *int `json:"doc_stride,omitempty"`
// The maximum length of predicted answers (e.g., only answers with a shorter
// length are considered).
MaxAnswerLen *int `json:"max_answer_len,omitempty"`
// The maximum length of the total sentence (context + question) in tokens of
// each chunk passed to the model. The context will be split in several chunks
// (using doc_stride as overlap) if needed.
MaxSeqLen *int `json:"max_seq_len,omitempty"`
// The maximum length of the question after tokenization. It will be truncated
// if needed.
MaxQuestionLen *int `json:"max_question_len,omitempty"`
// Whether to accept impossible as an answer.
HandleImpossibleAnswer *bool `json:"handle_impossible_answer,omitempty"`
// Attempts to align the answer to real words. Improves quality on space
// separated languages. Might hurt on non-space-separated languages (like
// Japanese or Chinese).
AlignToWords *bool `json:"align_to_words,omitempty"`
}
QuestionAnsweringParameters specify additional inference parameters for question answering.
func (*QuestionAnsweringParameters) Clone ¶
func (p *QuestionAnsweringParameters) Clone() QuestionAnsweringParameters
Clone returns a deep defensive copy of the parameters.
type QuestionAnsweringRequest ¶
type QuestionAnsweringRequest struct {
// The question and context pair to answer.
// Required.
Input QuestionAnsweringInput `json:"inputs"`
// Additional inference parameters for question answering.
Parameters *QuestionAnsweringParameters `json:"parameters,omitempty"`
}
QuestionAnsweringRequest represents a question answering inference request to the API.
func (*QuestionAnsweringRequest) Clone ¶
func (r *QuestionAnsweringRequest) Clone() QuestionAnsweringRequest
Clone returns a deep defensive copy of the request.
type RawService ¶
type RawService struct {
// contains filtered or unexported fields
}
RawService sends raw HTTP requests using the configured request options.
It is the deliberate exception to the rest of the SDK, where endpoints are exposed as Client methods: RawService is the advanced escape hatch for endpoints the SDK does not model type-safely. It combines several axes (byte-slice or io.Reader bodies, typed error handling or raw HTTP responses, one-shot or SSE streaming) into eight methods, which is enough surface that keeping it namespaced under Client.Raw() avoids cluttering the Client API.
RawService is immutable and holds a snapshot of the client options.
func (RawService) Do ¶
func (r RawService) Do( requestBody []byte, method string, path string, opts ...Option, ) (*http.Response, error)
Do performs a raw HTTP request with a byte slice body and applies SDK error interpretation on non-2xx responses. The caller must close resp.Body on success.
func (RawService) DoRaw ¶
func (r RawService) DoRaw( requestBody []byte, method string, path string, opts ...Option, ) (*http.Response, error)
DoRaw performs a raw HTTP request with a byte slice body without translating non-2xx responses into SDK errors. The caller must close resp.Body on success.
func (RawService) DoRawReader ¶
func (r RawService) DoRawReader( requestBody io.Reader, method string, path string, opts ...Option, ) (*http.Response, error)
DoRawReader performs a raw HTTP request with a streaming body without translating non-2xx responses into SDK errors. The caller must close resp.Body on success.
func (RawService) DoReader ¶
func (r RawService) DoReader( requestBody io.Reader, method string, path string, opts ...Option, ) (*http.Response, error)
DoReader performs a raw HTTP request with a streaming body and applies SDK error interpretation on non-2xx responses. The caller must close resp.Body on success.
func (RawService) Stream ¶
func (r RawService) Stream( requestBody []byte, method string, path string, opts ...Option, ) (*RawStream, error)
Stream performs a raw HTTP request and returns an SSE stream, applying SDK error interpretation on non-2xx responses. Callers should Close the returned RawStream when finished to promptly release the HTTP connection and decoder goroutine.
func (RawService) StreamRaw ¶
func (r RawService) StreamRaw( requestBody []byte, method string, path string, opts ...Option, ) (*RawStream, error)
StreamRaw performs a raw HTTP request and returns an SSE stream without translating non-2xx responses into SDK errors. This function is probably only interesting to advanced users. Only use this when you need to inspect the raw response; callers are responsible for interpreting HTTP errors themselves. Callers should Close the returned RawStream when finished to promptly release the HTTP connection and decoder goroutine.
func (RawService) StreamRawReader ¶
func (r RawService) StreamRawReader( requestBody io.Reader, method string, path string, opts ...Option, ) (*RawStream, error)
StreamRawReader performs a raw HTTP request with a streaming body and returns an SSE stream without translating non-2xx responses into SDK errors. This function is probably only interesting to advanced users. Only use this when you need to inspect the raw response; callers are responsible for interpreting HTTP errors themselves. Callers should Close the returned RawStream when finished to promptly release the HTTP connection and decoder goroutine.
func (RawService) StreamReader ¶
func (r RawService) StreamReader( requestBody io.Reader, method string, path string, opts ...Option, ) (*RawStream, error)
StreamReader performs a raw HTTP request with a streaming body and returns an SSE stream with SDK error interpretation. Callers should Close the returned RawStream when finished to promptly release the HTTP connection and decoder goroutine.
type RawStream ¶
type RawStream struct {
// contains filtered or unexported fields
}
RawStream exposes a raw SSE stream returned by RawService stream methods.
type ResponseFormatType ¶
type ResponseFormatType string
ResponseFormatType enumerates known response formats.
const ( // ResponseFormatTypeText requests a text response. ResponseFormatTypeText ResponseFormatType = "text" // ResponseFormatTypeJSONSchema requests a JSON schema response. ResponseFormatTypeJSONSchema ResponseFormatType = "json_schema" // ResponseFormatTypeJSONObject requests a JSON object response. ResponseFormatTypeJSONObject ResponseFormatType = "json_object" )
type SDKError ¶
SDKError represents a client-side SDK error that occurred before a response was received from the API.
Users can type-assert errors to *SDKError to access the error kind and underlying cause:
if sdkErr, ok := err.(*hfgo.SDKError); ok {
fmt.Printf("Kind %s: %s\n", sdkErr.Kind, sdkErr.Message)
}
type SDKErrorKind ¶
type SDKErrorKind = hferrors.SDKErrorKind
SDKErrorKind represents the category of a client-side SDK error.
type Summarization ¶
type Summarization struct {
// The summarized text.
SummaryText string `json:"summary_text"`
}
Summarization represents a summarization output.
type SummarizationBatchRequest ¶
type SummarizationBatchRequest struct {
// The texts to summarize.
// Required.
Inputs []string `json:"inputs"`
// Additional inference parameters for summarization.
Parameters *SummarizationParameters `json:"parameters,omitempty"`
}
SummarizationBatchRequest represents a batched summarization inference request to the API for multiple inputs.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.
func (*SummarizationBatchRequest) Clone ¶
func (r *SummarizationBatchRequest) Clone() SummarizationBatchRequest
Clone returns a deep defensive copy of the request.
type SummarizationParameters ¶
type SummarizationParameters struct {
// Whether to clean up the potential extra spaces in the text output.
CleanUpTokenizationSpaces *bool `json:"clean_up_tokenization_spaces,omitempty"`
// The truncation strategy to use.
Truncation *string `json:"truncation,omitempty"`
// GenerateParameters provides additional parametrization of the text
// generation algorithm (e.g. "max_new_tokens", "temperature", "top_k").
//
// It is exposed because it is part of the upstream inference schema, but the
// set of valid arguments is not documented by Hugging Face and may depend on
// the model being used. Invalid or unsupported arguments can be rejected by
// the API.
GenerateParameters map[string]any `json:"generate_parameters,omitempty"`
}
SummarizationParameters specify additional inference parameters for summarization tasks.
func (*SummarizationParameters) Clone ¶
func (p *SummarizationParameters) Clone() SummarizationParameters
Clone returns a deep defensive copy of the parameters. The GenerateParameters map is copied as a new map, but its values are shared because their types are not known statically.
type SummarizationRequest ¶
type SummarizationRequest struct {
// The text to summarize.
// Required.
Input string `json:"inputs"`
// Additional inference parameters for summarization.
Parameters *SummarizationParameters `json:"parameters,omitempty"`
}
SummarizationRequest represents a summarization inference request to the API for a single input.
func (*SummarizationRequest) Clone ¶
func (r *SummarizationRequest) Clone() SummarizationRequest
Clone returns a deep defensive copy of the request.
type TableQuestionAnswer ¶
type TableQuestionAnswer struct {
// The answer to the question given the table. If there is an aggregator,
// the answer will be preceded by "AGGREGATOR >".
Answer string `json:"answer"`
// Coordinates of the cells of the answers.
Coordinates [][]int `json:"coordinates"`
// List of strings made up of the answer cell values.
Cells []string `json:"cells"`
// If the model has an aggregator, this returns the aggregator.
Aggregator *string `json:"aggregator,omitempty"`
}
TableQuestionAnswer represents a table question answering output.
type TableQuestionAnsweringInput ¶
type TableQuestionAnsweringInput struct {
// The question to be answered about the table.
// Required.
Question string `json:"question"`
// The table to serve as context for the questions.
// Each key is a column name, and the value is a list of cell values for that column.
// Required.
Table map[string][]string `json:"table"`
}
TableQuestionAnsweringInput represents the input data for a table question answering request. Both Question and Table are required.
type TableQuestionAnsweringParameters ¶
type TableQuestionAnsweringParameters struct {
// Activates and controls padding.
Padding *string `json:"padding,omitempty"`
// Whether to do inference sequentially or as a batch. Batching is faster,
// but models like SQA require the inference to be done sequentially to
// extract relations within sequences, given their conversational nature.
Sequential *bool `json:"sequential,omitempty"`
// Activates and controls truncation.
Truncation *bool `json:"truncation,omitempty"`
}
TableQuestionAnsweringParameters specify additional inference parameters for table question answering tasks.
func (*TableQuestionAnsweringParameters) Clone ¶
func (p *TableQuestionAnsweringParameters) Clone() TableQuestionAnsweringParameters
Clone returns a deep defensive copy of the parameters.
type TableQuestionAnsweringRequest ¶
type TableQuestionAnsweringRequest struct {
// The question and table pair to answer.
// Required.
Input TableQuestionAnsweringInput `json:"inputs"`
// Additional inference parameters for table question answering.
Parameters *TableQuestionAnsweringParameters `json:"parameters,omitempty"`
}
TableQuestionAnsweringRequest represents a table question answering inference request to the API.
func (*TableQuestionAnsweringRequest) Clone ¶
func (r *TableQuestionAnsweringRequest) Clone() TableQuestionAnsweringRequest
Clone returns a deep defensive copy of the request.
type TextClassification ¶
type TextClassification struct {
// The predicted class label.
Label string `json:"label"`
// The corresponding probability.
Score float64 `json:"score"`
}
TextClassification represents a text classification output.
type TextClassificationBatchRequest ¶
type TextClassificationBatchRequest struct {
// The texts to classify.
// Required.
Inputs []string `json:"inputs"`
// Additional inference parameters for text classification
Parameters *TextClassificationParameters `json:"parameters,omitempty"`
}
TextClassificationBatchRequest represents a batched text classification inference request to the API for multiple inputs.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.
func (*TextClassificationBatchRequest) Clone ¶
func (r *TextClassificationBatchRequest) Clone() TextClassificationBatchRequest
Clone returns a deep defensive copy of the request.
type TextClassificationParameters ¶
type TextClassificationParameters struct {
// Possible values: sigmoid, softmax, none.
FunctionToApply *string `json:"function_to_apply,omitempty"`
// When specified, limits the output to the top K most probable classes.
TopK *int `json:"top_k,omitempty"`
}
TextClassificationParameters specify additional inference parameters for text classification.
func (*TextClassificationParameters) Clone ¶
func (p *TextClassificationParameters) Clone() TextClassificationParameters
Clone returns a deep defensive copy of the parameters.
type TextClassificationRequest ¶
type TextClassificationRequest struct {
// The text to classify.
// Required.
Input string `json:"inputs"`
// Additional inference parameters for text classification
Parameters *TextClassificationParameters `json:"parameters,omitempty"`
}
TextClassificationRequest represents a text classification inference request to the API for a single input.
func (*TextClassificationRequest) Clone ¶
func (r *TextClassificationRequest) Clone() TextClassificationRequest
Clone returns a deep defensive copy of the request.
type TokenClassification ¶
type TokenClassification struct {
// The predicted label for a group of one or more tokens.
// Present when an aggregation strategy other than "none" is used.
EntityGroup *string `json:"entity_group,omitempty"`
// The predicted label for a single token.
// Present when aggregation_strategy is "none".
Entity *string `json:"entity,omitempty"`
// The associated score / probability.
Score float64 `json:"score"`
// The corresponding text.
Word string `json:"word"`
// The character position in the input where this group begins.
Start int `json:"start"`
// The character position in the input where this group ends.
End int `json:"end"`
}
TokenClassification represents a token classification output.
type TokenClassificationBatchRequest ¶
type TokenClassificationBatchRequest struct {
// The texts to classify tokens from.
// Required.
Inputs []string `json:"inputs"`
// Additional inference parameters for token classification.
Parameters *TokenClassificationParameters `json:"parameters,omitempty"`
}
TokenClassificationBatchRequest represents a batched token classification request to the API for multiple inputs.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.
func (*TokenClassificationBatchRequest) Clone ¶
func (r *TokenClassificationBatchRequest) Clone() TokenClassificationBatchRequest
Clone returns a deep defensive copy of the request.
type TokenClassificationParameters ¶
type TokenClassificationParameters struct {
// A list of labels to ignore in the classification results.
IgnoreLabels []string `json:"ignore_labels,omitempty"`
// The number of overlapping tokens between chunks when splitting the input text.
Stride *int `json:"stride,omitempty"`
// The strategy used to fuse tokens based on model predictions.
// Possible values: "none", "simple", "first", "average", "max".
AggregationStrategy *string `json:"aggregation_strategy,omitempty"`
}
TokenClassificationParameters specify additional inference parameters for token classification.
func (*TokenClassificationParameters) Clone ¶
func (p *TokenClassificationParameters) Clone() TokenClassificationParameters
Clone returns a deep defensive copy of the parameters.
type TokenClassificationRequest ¶
type TokenClassificationRequest struct {
// The text to classify tokens from.
// Required.
Input string `json:"inputs"`
// Additional inference parameters for token classification.
Parameters *TokenClassificationParameters `json:"parameters,omitempty"`
}
TokenClassificationRequest represents a token classification inference request to the API for a single input.
func (*TokenClassificationRequest) Clone ¶
func (r *TokenClassificationRequest) Clone() TokenClassificationRequest
Clone returns a deep defensive copy of the request.
type ToolChoiceMode ¶
type ToolChoiceMode string
ToolChoiceMode enumerates known tool choice modes.
const ( // ToolChoiceModeAuto lets the provider decide the tool choice. ToolChoiceModeAuto ToolChoiceMode = "auto" // ToolChoiceModeNone disables tool usage. ToolChoiceModeNone ToolChoiceMode = "none" // ToolChoiceModeRequired requires tool usage. ToolChoiceModeRequired ToolChoiceMode = "required" )
type Translation ¶
type Translation struct {
// The translated text.
TranslationText string `json:"translation_text"`
}
Translation represents a translation output.
type TranslationBatchRequest ¶
type TranslationBatchRequest struct {
// The texts to translate.
// Required.
Inputs []string `json:"inputs"`
// Additional inference parameters for translation.
Parameters *TranslationParameters `json:"parameters,omitempty"`
}
TranslationBatchRequest represents a batched translation inference request to the API for multiple inputs.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.
func (*TranslationBatchRequest) Clone ¶
func (r *TranslationBatchRequest) Clone() TranslationBatchRequest
Clone returns a deep defensive copy of the request.
type TranslationParameters ¶
type TranslationParameters struct {
// Whether to clean up the potential extra spaces in the text output.
CleanUpTokenizationSpaces *bool `json:"clean_up_tokenization_spaces,omitempty"`
// The source language of the text. Required for models that can
// translate from multiple languages.
SrcLang *string `json:"src_lang,omitempty"`
// The target language to translate to. Required for models that can
// translate to multiple languages.
TgtLang *string `json:"tgt_lang,omitempty"`
// The truncation strategy to use.
Truncation *string `json:"truncation,omitempty"`
// GenerateParameters provides additional parametrization of the text
// generation algorithm (e.g. "max_new_tokens", "temperature", "top_k").
//
// It is exposed because it is part of the upstream inference schema, but the
// set of valid arguments is not documented by Hugging Face and may depend on
// the model being used. Invalid or unsupported arguments can be rejected by
// the API.
GenerateParameters map[string]any `json:"generate_parameters,omitempty"`
}
TranslationParameters specify additional inference parameters for translation tasks.
func (*TranslationParameters) Clone ¶
func (p *TranslationParameters) Clone() TranslationParameters
Clone returns a deep defensive copy of the parameters. The GenerateParameters map is copied as a new map, but its values are shared because their types are not known statically.
type TranslationRequest ¶
type TranslationRequest struct {
// The text to translate.
// Required.
Input string `json:"inputs"`
// Additional inference parameters for translation.
Parameters *TranslationParameters `json:"parameters,omitempty"`
}
TranslationRequest represents a translation inference request to the API for a single input.
func (*TranslationRequest) Clone ¶
func (r *TranslationRequest) Clone() TranslationRequest
Clone returns a deep defensive copy of the request.
type ZeroShotTextClassification ¶
type ZeroShotTextClassification struct {
// The predicted class label.
Label string `json:"label"`
// The corresponding probability.
Score float64 `json:"score"`
}
ZeroShotTextClassification represents a zero-shot text classification output.
type ZeroShotTextClassificationBatchRequest ¶
type ZeroShotTextClassificationBatchRequest struct {
// The texts to classify.
// Required.
Inputs []string `json:"inputs"`
// Additional inference parameters for zero-shot text classification.
// Required.
Parameters *ZeroShotTextClassificationParameters `json:"parameters,omitempty"`
}
ZeroShotTextClassificationBatchRequest represents a batched zero-shot text classification request to the API for multiple inputs.
NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.
func (*ZeroShotTextClassificationBatchRequest) Clone ¶
func (r *ZeroShotTextClassificationBatchRequest) Clone() ZeroShotTextClassificationBatchRequest
Clone returns a deep defensive copy of the request.
type ZeroShotTextClassificationParameters ¶
type ZeroShotTextClassificationParameters struct {
// The set of possible class labels to classify the text into.
// Required.
CandidateLabels []string `json:"candidate_labels,omitempty"`
// The sentence used in conjunction with candidate_labels to attempt
// the text classification by replacing the placeholder with the
// candidate labels.
HypothesisTemplate *string `json:"hypothesis_template,omitempty"`
// Whether multiple candidate labels can be true. If false, the scores
// are normalized such that the sum of the label likelihoods for each
// sequence is 1. If true, the labels are considered independent and
// probabilities are normalized for each candidate.
MultiLabel *bool `json:"multi_label,omitempty"`
}
ZeroShotTextClassificationParameters specify additional inference parameters for zero-shot text classification.
func (*ZeroShotTextClassificationParameters) Clone ¶
func (p *ZeroShotTextClassificationParameters) Clone() ZeroShotTextClassificationParameters
Clone returns a deep defensive copy of the parameters.
type ZeroShotTextClassificationRequest ¶
type ZeroShotTextClassificationRequest struct {
// The text to classify.
// Required.
Input string `json:"inputs"`
// Additional inference parameters for zero-shot text classification.
// Required.
Parameters *ZeroShotTextClassificationParameters `json:"parameters,omitempty"`
}
ZeroShotTextClassificationRequest represents a zero-shot text classification request to the API for a single input.
func (*ZeroShotTextClassificationRequest) Clone ¶
func (r *ZeroShotTextClassificationRequest) Clone() ZeroShotTextClassificationRequest
Clone returns a deep defensive copy of the request.
Source Files
¶
- chat_common.go
- chat_request.go
- chat_response.go
- chat_service.go
- chat_streaming.go
- client.go
- clone.go
- doc.go
- errors.go
- fill_mask.go
- fill_mask_service.go
- model_request.go
- option.go
- options.go
- question_answering.go
- question_answering_service.go
- raw_service.go
- summarization.go
- summarization_service.go
- table_question_answering.go
- table_question_answering_service.go
- text_classification.go
- text_classification_service.go
- token_classification.go
- token_classification_service.go
- translation.go
- translation_service.go
- version.go
- zero_shot_text_classification.go
- zero_shot_text_classification_service.go
Directories
¶
| Path | Synopsis |
|---|---|
|
examples
|
|
|
chat/basic
command
|
|
|
chat/convo
command
|
|
|
chat/streaming
command
|
|
|
fill-mask/basic
command
|
|
|
fill-mask/batch
command
|
|
|
fill-mask/params
command
|
|
|
question-answering/basic
command
|
|
|
question-answering/params
command
|
|
|
summarization/basic
command
|
|
|
summarization/batch
command
|
|
|
summarization/params
command
|
|
|
table-question-answering/basic
command
|
|
|
table-question-answering/params
command
|
|
|
text-classification/basic
command
|
|
|
text-classification/batch
command
|
|
|
token-classification/basic
command
|
|
|
token-classification/batch
command
|
|
|
token-classification/params
command
|
|
|
translation/basic
command
|
|
|
translation/batch
command
|
|
|
translation/params
command
|
|
|
internal
|
|
|
chatstream
Package chatstream provides helpers for working with streamed chat responses.
|
Package chatstream provides helpers for working with streamed chat responses. |
|
hferrors
Package hferrors defines the reusable SDK error types returned by hfgo.
|
Package hferrors defines the reusable SDK error types returned by hfgo. |
|
request
Package request contains the lower-level HTTP, SSE, and JSON utilities used by the SDK to build, send, and process Hugging Face API requests.
|
Package request contains the lower-level HTTP, SSE, and JSON utilities used by the SDK to build, send, and process Hugging Face API requests. |
|
sdkversion
Package sdkversion exposes the SDK version string used in User-Agent headers.
|
Package sdkversion exposes the SDK version string used in User-Agent headers. |
|
testutils
Package testutils provides helpers for tests.
|
Package testutils provides helpers for tests. |