hfgo

package module
v4.0.0-rc1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 29, 2026 License: MIT Imports: 14 Imported by: 0

README

hfgo

Build Unit Tests Integration Tests Lint CodeQL Go Reference OpenSSF Scorecard OpenSSF Best Practices

An unofficial Go SDK for the Hugging Face Inference API. Directly call any model available in the Model Hub.

An API key is required for authorized access. To get one, create a Hugging Face account and generate a token.

⚠️ v4 Release Candidate

v4 is currently in release candidate status (v4.0.0-rc1). The API may evolve before the final v4.0.0 release. v3 and earlier are deprecated and no longer maintained. See #72 for more information.

Usage

package main

import (
	"fmt"
	"log"
	"os"

	"github.com/Kardbord/hfgo/v4"
)

func main() {
	token := os.Getenv("HF_TOKEN")
	if token == "" {
		log.Fatal("HF_TOKEN environment variable is not set")
	}

	client := hfgo.NewClient(
		hfgo.WithToken(token),
		hfgo.WithModel("deepseek-ai/DeepSeek-R1"),
	)

	request := hfgo.ChatRequest{
		Messages: []hfgo.ChatMessage{
			{
				Role: "user",
				Content: hfgo.ChatMessageContent{
					Text: Ptr("Hello! What is the capital of France?"),
				},
			},
		},
		MaxTokens: Ptr(1024),
	}

	response, err := client.Chat(request)
	if err != nil {
		log.Fatalf("Failed to complete chat request: %v", err)
	}

	for _, choice := range response.Choices {
		if choice.Message.Content != nil {
			fmt.Println(*choice.Message.Content)
		}
	}
}

func Ptr[T any](v T) *T {
	return &v
}

See the examples directory for more.

Inference Tasks

Contributing

Contributions are welcome in many forms!

  • Opening or commenting on issues to suggest new features, clarify requirements, and report bugs
  • Reviewing PRs to help improve code quality
  • Documentation improvements (updating README, docs/, or examples/)
  • Community engagement (helping new contributors, answering questions)

If you plan to contribute code, please open an issue first to discuss your proposed changes, coordinate with maintainers, and avoid duplicate work.

See CONTRIBUTING.md for detailed guidelines.

Resources

Documentation

Overview

Package hfgo provides Go bindings for the Hugging Face Inference API.

Design notes:

  • Clients are immutable; options are fixed at creation time, and each call snapshots them.
  • Clients are safe for concurrent use by default or when configured with immutable or synchronized dependencies.
  • Per-request options can override client defaults for a single call.
  • Request options are applied by value with defensive header copies; contexts and HTTP clients are shared.
  • HTTP client injection uses a value factory; return a fresh client value to avoid shared state.
  • The SDK favors upstream feature parity and uses DTOs closely aligned to the API; breaking changes are possible as the upstream API evolves.
  • WithDefaultHTTPClient restores the default client; a nil factory is treated as a configuration error.
  • Client.Raw() returns the RawService escape hatch for arbitrary endpoints, exposing both error-interpreting and raw request paths (Do vs DoRaw).
  • DTO validation is enforced during JSON marshal/unmarshal. Invalid request payloads surface as configuration errors. For responses, invalid content type surfaces as validation errors, while malformed JSON surfaces as serialization errors.
  • Concurrency assumes externally supplied objects (for example, transports) are not mutated after use unless callers provide their own synchronization.
  • Request DTOs are passed to Client methods by value and the SDK never mutates the caller's payload.
  • A request and the nested data it references must be treated as read-only while a call is in flight. For concurrent invocation, pass a defensive copy per call (for example, go client.Chat(req.Clone(), ...)) or build a fresh request per call. Every request DTO provides a deep Clone method.

Index

Constants

View Source
const (
	// SDKErrorKindValidation indicates a validation error in API responses.
	SDKErrorKindValidation = hferrors.SDKErrorKindValidation
	// SDKErrorKindConfiguration indicates invalid or missing configuration.
	SDKErrorKindConfiguration = hferrors.SDKErrorKindConfiguration
	// SDKErrorKindSerialization indicates a serialization or deserialization error.
	SDKErrorKindSerialization = hferrors.SDKErrorKindSerialization
	// SDKErrorKindTransport indicates a transport-layer failure.
	SDKErrorKindTransport = hferrors.SDKErrorKindTransport
	// SDKErrorKindInternal indicates an internal SDK error.
	SDKErrorKindInternal = hferrors.SDKErrorKindInternal
)
View Source
const (
	// SummarizationTruncationDoNotTruncate keeps the input as-is without any truncation.
	SummarizationTruncationDoNotTruncate = "do_not_truncate"
	// SummarizationTruncationLongestFirst truncates the longest side first when the input exceeds the model's maximum length.
	SummarizationTruncationLongestFirst = "longest_first"
	// SummarizationTruncationOnlyFirst trims the input on the first (left) side when it exceeds the model's maximum length.
	SummarizationTruncationOnlyFirst = "only_first"
	// SummarizationTruncationOnlySecond trims the input on the second (right) side when it exceeds the model's maximum length.
	SummarizationTruncationOnlySecond = "only_second"
)
View Source
const (
	// TableQuestionAnsweringPaddingDoNotPad does not pad the input.
	TableQuestionAnsweringPaddingDoNotPad = "do_not_pad"
	// TableQuestionAnsweringPaddingLongest pads to the longest sequence in the batch.
	TableQuestionAnsweringPaddingLongest = "longest"
	// TableQuestionAnsweringPaddingMaxLength pads to the maximum length.
	TableQuestionAnsweringPaddingMaxLength = "max_length"
)
View Source
const (
	// TextClassificationFuncSigmoid applies a sigmoid to each score independently.
	// Useful for multi-label classification tasks, where multiple classes may apply simultaneously.
	TextClassificationFuncSigmoid = "sigmoid"
	// TextClassificationFuncSoftmax normalizes scores into a probability distribution summing to 1.
	// Useful for single-label multi-class classification tasks, where exactly one class applies.
	TextClassificationFuncSoftmax = "softmax"
	// TextClassificationFuncNone returns raw scores without any transformation applied.
	TextClassificationFuncNone = "none"
)
View Source
const (
	// TokenClassificationAggregationNone does not aggregate tokens.
	// Each token is classified individually.
	TokenClassificationAggregationNone = "none"
	// TokenClassificationAggregationSimple groups consecutive tokens with the same
	// label into a single entity.
	TokenClassificationAggregationSimple = "simple"
	// TokenClassificationAggregationFirst is similar to "simple", but also preserves
	// word integrity using the label predicted for the first token in a word.
	TokenClassificationAggregationFirst = "first"
	// TokenClassificationAggregationAverage is similar to "simple", but also preserves
	// word integrity using the label with the highest score, averaged across the
	// word's tokens.
	TokenClassificationAggregationAverage = "average"
	// TokenClassificationAggregationMax is similar to "simple", but also preserves
	// word integrity using the label with the highest score across the word's tokens.
	TokenClassificationAggregationMax = "max"
)
View Source
const (
	// TranslationTruncationDoNotTruncate keeps the input as-is without any truncation.
	TranslationTruncationDoNotTruncate = "do_not_truncate"
	// TranslationTruncationLongestFirst truncates the longest side first when the input exceeds the model's maximum length.
	TranslationTruncationLongestFirst = "longest_first"
	// TranslationTruncationOnlyFirst trims the input on the first (left) side when it exceeds the model's maximum length.
	TranslationTruncationOnlyFirst = "only_first"
	// TranslationTruncationOnlySecond trims the input on the second (right) side when it exceeds the model's maximum length.
	TranslationTruncationOnlySecond = "only_second"
)
View Source
const EndpointChatCompletion = "/v1/chat/completions"

EndpointChatCompletion specifies the chat completion endpoint.

View Source
const Version = sdkversion.Version

Version is the current version of the hfgo SDK. This follows semantic versioning (semver.org).

Variables

This section is empty.

Functions

func UserAgent

func UserAgent() string

UserAgent returns the User-Agent string used for HTTP requests. The format is: hfgo/<version> (Go).

Types

type APIError

type APIError = hferrors.APIError

APIError represents an error returned by the HuggingFace API. It includes the HTTP status code, error message, response body, and request ID if available.

Users can type-assert errors to *APIError to access additional error information and helper methods:

if apiErr, ok := err.(*hfgo.APIError); ok {
    if apiErr.IsAuthenticationError() {
        // Handle authentication error
    }
}

type ChatChoice

type ChatChoice struct {
	// Required.
	FinishReason string `json:"finish_reason"`
	// Required.
	Index    int           `json:"index"`
	LogProbs *ChatLogProbs `json:"logprobs,omitempty"`
	// Required.
	Message ChatCompletionMessage `json:"message"`
}

ChatChoice is a single non-streaming completion choice.

type ChatCompletionMessage

type ChatCompletionMessage struct {
	Role string `json:"role"`
	// Content is present for text responses.
	Content *string `json:"content,omitempty"`
	// ToolCallID is set when returning tool-specific content.
	ToolCallID *string `json:"tool_call_id,omitempty"`
	// ToolCalls is present for tool call responses.
	ToolCalls []ChatToolCallOutput `json:"tool_calls,omitempty"`
}

ChatCompletionMessage is a message returned by the model. It is either a text message (Content) or a tool call message (ToolCalls).

func (*ChatCompletionMessage) UnmarshalJSON

func (m *ChatCompletionMessage) UnmarshalJSON(data []byte) error

UnmarshalJSON enforces the union shape for ChatCompletionMessage.

type ChatFunctionCall

type ChatFunctionCall struct {
	// Required.
	Name string `json:"name"`
	// Required.
	Arguments   string  `json:"arguments"`
	Description *string `json:"description,omitempty"`
}

ChatFunctionCall represents a tool function call with arguments.

func (*ChatFunctionCall) Clone

Clone returns a deep defensive copy of the function call.

func (ChatFunctionCall) MarshalJSON

func (f ChatFunctionCall) MarshalJSON() ([]byte, error)

MarshalJSON enforces required fields on ChatFunctionCall.

func (*ChatFunctionCall) UnmarshalJSON

func (f *ChatFunctionCall) UnmarshalJSON(data []byte) error

UnmarshalJSON enforces required fields on ChatFunctionCall.

type ChatFunctionDefinition

type ChatFunctionDefinition struct {
	// Required.
	Name string `json:"name"`
	// A description of what the function does.
	Description *string `json:"description,omitempty"`
	// JSON schema describing function parameters.
	Parameters json.RawMessage `json:"parameters,omitempty"`
}

ChatFunctionDefinition describes a callable function.

func (*ChatFunctionDefinition) Clone

Clone returns a deep defensive copy of the function definition.

func (ChatFunctionDefinition) MarshalJSON

func (f ChatFunctionDefinition) MarshalJSON() ([]byte, error)

MarshalJSON enforces required fields on ChatFunctionDefinition.

type ChatFunctionName

type ChatFunctionName struct {
	// Required.
	Name string `json:"name"`
}

ChatFunctionName identifies a tool function by name.

func (*ChatFunctionName) Clone

Clone returns a defensive copy of the function name.

func (ChatFunctionName) MarshalJSON

func (f ChatFunctionName) MarshalJSON() ([]byte, error)

MarshalJSON enforces required fields on ChatFunctionName.

type ChatImageURL

type ChatImageURL struct {
	// Required.
	URL string `json:"url"`
}

ChatImageURL contains the URL for an image chunk.

func (*ChatImageURL) Clone

func (u *ChatImageURL) Clone() ChatImageURL

Clone returns a defensive copy of the image URL.

func (ChatImageURL) MarshalJSON

func (u ChatImageURL) MarshalJSON() ([]byte, error)

MarshalJSON enforces the required URL on ChatImageURL.

type ChatJSONSchemaConfig

type ChatJSONSchemaConfig struct {
	// The name of the response format.
	// Required.
	Name string `json:"name"`
	// A description of what the response format is for.
	Description *string `json:"description,omitempty"`
	// The schema for the response format as a JSON Schema object.
	// Learn how to build JSON schemas at https://json-schema.org/.
	Schema json.RawMessage `json:"schema,omitempty"`
	// Whether to enable strict schema adherence.
	Strict *bool `json:"strict,omitempty"`
}

ChatJSONSchemaConfig defines JSON schema response formatting.

func (*ChatJSONSchemaConfig) Clone

Clone returns a deep defensive copy of the JSON schema configuration.

type ChatLogProb

type ChatLogProb struct {
	// Required.
	Token string `json:"token"`
	// Required.
	LogProb float64 `json:"logprob"`
	// Required.
	TopLogProbs []ChatTopLogProb `json:"top_logprobs"`
}

ChatLogProb contains logprob information for a token.

type ChatLogProbs

type ChatLogProbs struct {
	// Required.
	Content []ChatLogProb `json:"content"`
}

ChatLogProbs contains per-token log probabilities.

type ChatMessage

type ChatMessage struct {
	// Role of the message author (for example: system, user, assistant, tool).
	// Required.
	Role string `json:"role"`
	// Optional name for the participant.
	Name *string `json:"name,omitempty"`

	// Content may be a string or a list of content chunks.
	// Either Content or ToolCalls should be supplied.
	Content ChatMessageContent `json:"content"`

	// ToolCalls is used instead of Content when providing tool call messages.
	ToolCalls []ChatToolCall `json:"tool_calls,omitempty"`
}

ChatMessage represents a single chat message.

func (*ChatMessage) Clone

func (m *ChatMessage) Clone() ChatMessage

Clone returns a deep defensive copy of the message.

func (ChatMessage) MarshalJSON

func (m ChatMessage) MarshalJSON() ([]byte, error)

MarshalJSON enforces the union shape for ChatMessage.

type ChatMessageChunk

type ChatMessageChunk struct {
	// Required when Type is "text".
	Text *string `json:"text,omitempty"`
	// Required when Type is "image_url".
	ImageURL *ChatImageURL `json:"image_url,omitempty"`
	// Required. Possible values: text, image_url.
	Type MessageChunkType `json:"type"`
}

ChatMessageChunk represents a content chunk (text or image URL). Type must be "text" with Text set, or "image_url" with ImageURL set.

func (*ChatMessageChunk) Clone

Clone returns a deep defensive copy of the chunk.

func (ChatMessageChunk) MarshalJSON

func (c ChatMessageChunk) MarshalJSON() ([]byte, error)

MarshalJSON enforces the union shape for ChatMessageChunk.

type ChatMessageContent

type ChatMessageContent struct {

	// Text holds plain string content.
	Text *string `json:"-"`
	// Chunks holds structured content chunks.
	Chunks []ChatMessageChunk `json:"-"`
}

ChatMessageContent can be a string or []ChatMessageChunk. Use a string for pure text content or a chunk list for multimodal content.

func (*ChatMessageContent) Clone

Clone returns a deep defensive copy of the content.

func (ChatMessageContent) MarshalJSON

func (c ChatMessageContent) MarshalJSON() ([]byte, error)

MarshalJSON enforces the union shape for ChatMessageContent.

type ChatRequest

type ChatRequest struct {
	// Model to use for the chat completion.
	// Required.
	Model *string `json:"model,omitempty"`

	// Number between -2.0 and 2.0. Positive values penalize new tokens based
	// on their existing frequency in the text so far, decreasing the model's
	// likelihood to repeat the same line verbatim.
	FrequencyPenalty *float64 `json:"frequency_penalty,omitempty"`

	// Whether to return log probabilities of the output tokens or not.
	// If true, returns the log probabilities of each output token returned
	// in the content of message.
	LogProbs *bool `json:"logprobs,omitempty"`

	// The maximum number of tokens that can be generated in the chat completion.
	// Default: 1024. Minimum: 0.
	MaxTokens *int `json:"max_tokens,omitempty"`

	// A list of messages comprising the conversation so far.
	// Required.
	Messages []ChatMessage `json:"messages"`

	// Number between -2.0 and 2.0. Positive values penalize new tokens based
	// on whether they appear in the text so far, increasing the model's
	// likelihood to talk about new topics.
	PresencePenalty *float64 `json:"presence_penalty,omitempty"`

	// Response format configuration. Known types: text, json_schema, json_object.
	// Non-empty provider-specific values are also accepted.
	ResponseFormat *ChatResponseFormat `json:"response_format,omitempty"`

	// Seed for deterministic sampling.
	// Minimum: 0.
	Seed *int64 `json:"seed,omitempty"`

	// Up to 4 sequences where the API will stop generating further tokens.
	Stop []string `json:"stop,omitempty"`

	// If true, generated tokens are returned as a stream using SSE.
	// For more information about streaming, see:
	// https://huggingface.co/docs/text-generation-inference/conceptual/streaming
	Stream *bool `json:"stream,omitempty"`

	// Stream options for SSE responses.
	StreamOptions *ChatStreamOptions `json:"stream_options,omitempty"`

	// Sampling temperature to use, between 0 and 2.
	// We generally recommend altering this or TopP but not both.
	Temperature *float64 `json:"temperature,omitempty"`

	// Tool choice behavior. Known values: auto, none, required, or
	// {"function":{"name":"..."}}. Meanings:
	// - auto: model can pick between generating a message or calling tools.
	// - none: model will not call any tool and instead generates a message.
	// - required: model must call one or more tools.
	// Non-empty provider-specific values are also accepted.
	ToolChoice *ChatToolChoice `json:"tool_choice,omitempty"`

	// A prompt to be appended before the tools.
	ToolPrompt *string `json:"tool_prompt,omitempty"`

	// A list of tools the model may call.
	// Currently, only functions are supported as a tool. Use this to provide
	// a list of functions the model may generate JSON inputs for.
	Tools []ChatTool `json:"tools,omitempty"`

	// Number of most likely tokens to return at each token position.
	// LogProbs must be true if this is set.
	// An integer between 0 and 5.
	TopLogProbs *int `json:"top_logprobs,omitempty"`

	// Nucleus sampling probability mass.
	// For example, 0.1 means only the tokens comprising the top 10% probability
	// mass are considered.
	TopP *float64 `json:"top_p,omitempty"`
}

ChatRequest represents a completion request for the chat API. Output type depends on the Stream parameter.

func (*ChatRequest) Clone

func (r *ChatRequest) Clone() ChatRequest

Clone returns a deep defensive copy of the request. All reference fields — slices, maps, pointers, and JSON blobs — are copied so the returned value shares no backing storage with the receiver. This makes it safe to reuse a template request across concurrent calls, e.g. go client.Chat(req.Clone(), ...).

func (ChatRequest) MarshalJSON

func (r ChatRequest) MarshalJSON() ([]byte, error)

MarshalJSON enforces required fields for ChatRequest.

type ChatResponse

type ChatResponse struct {
	// Required.
	ID string `json:"id"`
	// Unix timestamp in seconds.
	// Required.
	Created int64 `json:"created"`
	// Required.
	Model string `json:"model"`
	// Required.
	SystemFingerprint string `json:"system_fingerprint"`
	// Required.
	Choices []ChatChoice `json:"choices"`
	// Required.
	Usage ChatUsage `json:"usage"`
}

ChatResponse represents a non-streaming completion response from the chat API. This is returned when Stream is false (the default).

type ChatResponseFormat

type ChatResponseFormat struct {
	// Known type values: text, json_schema, json_object.
	// Non-empty provider-specific values are also accepted.
	// For json_schema, JSONSchema is required.
	Type       ResponseFormatType    `json:"type"`
	JSONSchema *ChatJSONSchemaConfig `json:"json_schema,omitempty"`
}

ChatResponseFormat configures the response format.

func (*ChatResponseFormat) Clone

Clone returns a deep defensive copy of the response format.

func (ChatResponseFormat) MarshalJSON

func (r ChatResponseFormat) MarshalJSON() ([]byte, error)

MarshalJSON enforces the union shape for ChatResponseFormat.

type ChatStream

type ChatStream struct {
	// contains filtered or unexported fields
}

ChatStream wraps a streaming chat completion response.

func (*ChatStream) Close

func (c *ChatStream) Close() error

Close releases the underlying stream resources.

func (*ChatStream) Recv

Recv blocks until the next streaming chunk arrives or the context is done.

type ChatStreamChoice

type ChatStreamChoice struct {
	// Required.
	Delta        ChatStreamDelta `json:"delta"`
	FinishReason *string         `json:"finish_reason,omitempty"`
	// Required.
	Index    int           `json:"index"`
	LogProbs *ChatLogProbs `json:"logprobs,omitempty"`
}

ChatStreamChoice is a single streaming completion choice.

type ChatStreamDelta

type ChatStreamDelta struct {
	// Content is present for text deltas.
	Content *string `json:"content,omitempty"`
	// Role may be included with the first delta.
	Role *string `json:"role,omitempty"`
	// ToolCallID may be included for tool-specific content.
	ToolCallID *string `json:"tool_call_id,omitempty"`
	// ToolCalls is present for tool call deltas.
	ToolCalls []ChatStreamToolCall `json:"tool_calls,omitempty"`
}

ChatStreamDelta holds incremental updates for a stream. Deltas may include content/role/tool_call_id or role/tool_calls.

func (*ChatStreamDelta) UnmarshalJSON

func (d *ChatStreamDelta) UnmarshalJSON(data []byte) error

UnmarshalJSON enforces the union shape for ChatStreamDelta.

type ChatStreamFunction

type ChatStreamFunction struct {
	Name      string `json:"name"`
	Arguments string `json:"arguments"`
}

ChatStreamFunction represents a streamed function call.

type ChatStreamOptions

type ChatStreamOptions struct {
	// If set, an additional chunk is streamed before the [DONE] message
	// showing overall usage. The usage field on this chunk shows the token
	// usage statistics for the entire request, and the choices field will
	// always be an empty array. All other chunks include a usage field with
	// a null value.
	IncludeUsage *bool `json:"include_usage,omitempty"`
}

ChatStreamOptions configures streaming behavior.

func (*ChatStreamOptions) Clone

Clone returns a deep defensive copy of the stream options.

type ChatStreamResponse

type ChatStreamResponse struct {
	ID string `json:"id"`
	// Unix timestamp in seconds.
	Created           int64              `json:"created"`
	Model             string             `json:"model"`
	SystemFingerprint string             `json:"system_fingerprint"`
	Choices           []ChatStreamChoice `json:"choices"`
	Usage             *ChatUsage         `json:"usage,omitempty"`
}

ChatStreamResponse represents a streaming response chunk. This is returned when Stream is true.

type ChatStreamToolCall

type ChatStreamToolCall struct {
	ID       string             `json:"id,omitempty"`
	Type     string             `json:"type,omitempty"`
	Index    int                `json:"index"`
	Function ChatStreamFunction `json:"function"`
}

ChatStreamToolCall represents a tool call within a streaming delta.

type ChatTool

type ChatTool struct {
	// Tool type (currently only function is supported).
	Type     string                 `json:"type"`
	Function ChatFunctionDefinition `json:"function"`
}

ChatTool represents a tool definition provided to the model.

func (*ChatTool) Clone

func (t *ChatTool) Clone() ChatTool

Clone returns a deep defensive copy of the tool.

func (ChatTool) MarshalJSON

func (t ChatTool) MarshalJSON() ([]byte, error)

MarshalJSON enforces the tool shape for ChatTool.

type ChatToolCall

type ChatToolCall struct {
	// Tool call ID.
	// Required.
	ID string `json:"id"`
	// Tool call type (for example: function).
	// Required.
	Type string `json:"type"`
	// Required.
	Function ChatFunctionCall `json:"function"`
}

ChatToolCall represents a tool call in a message.

func (*ChatToolCall) Clone

func (c *ChatToolCall) Clone() ChatToolCall

Clone returns a deep defensive copy of the tool call.

func (ChatToolCall) MarshalJSON

func (c ChatToolCall) MarshalJSON() ([]byte, error)

MarshalJSON enforces the tool call shape for ChatToolCall.

type ChatToolCallOutput

type ChatToolCallOutput struct {
	// Required.
	ID string `json:"id"`
	// Required.
	Type string `json:"type"`
	// Required.
	Function ChatFunctionCall `json:"function"`
}

ChatToolCallOutput represents a tool call in a response message.

func (*ChatToolCallOutput) UnmarshalJSON

func (c *ChatToolCallOutput) UnmarshalJSON(data []byte) error

UnmarshalJSON enforces the tool call shape for ChatToolCallOutput.

type ChatToolChoice

type ChatToolChoice struct {

	// Mode is a tool choice mode. Known values: auto, none, required.
	// Non-empty provider-specific values are also accepted.
	Mode *ToolChoiceMode `json:"-"`
	// Function selects a specific tool function by name.
	Function *ChatFunctionName `json:"-"`
}

ChatToolChoice represents the tool choice union type. Either Mode is set to auto/none/required, or Function is set.

func (*ChatToolChoice) Clone

func (t *ChatToolChoice) Clone() ChatToolChoice

Clone returns a deep defensive copy of the tool choice.

func (ChatToolChoice) MarshalJSON

func (t ChatToolChoice) MarshalJSON() ([]byte, error)

MarshalJSON enforces the union shape for ChatToolChoice.

type ChatTopLogProb

type ChatTopLogProb struct {
	// Required.
	Token string `json:"token"`
	// Required.
	LogProb float64 `json:"logprob"`
}

ChatTopLogProb is a top log probability entry for a token position.

type ChatUsage

type ChatUsage struct {
	// Required.
	CompletionTokens int `json:"completion_tokens"`
	// Required.
	PromptTokens int `json:"prompt_tokens"`
	// Required.
	TotalTokens int `json:"total_tokens"`
}

ChatUsage contains token usage statistics.

type Client

type Client struct {
	// contains filtered or unexported fields
}

Client represents a HuggingFace API client with configured request options. Client instances are immutable; options are fixed at creation time and never mutated. This keeps client usage safe across goroutines and avoids surprises from mutable state. If options include externally-owned pointers, callers must avoid mutating them after creation or ensure their own synchronization. RawService captures a snapshot of these options when created.

func NewClient

func NewClient(opts ...Option) Client

NewClient creates a new Client instance with the provided request options. If no options are provided, default options will be used. Clients are immutable; to change options, create a new Client to keep calls deterministic.

func (Client) AnswerQuestion

func (c Client) AnswerQuestion(
	req QuestionAnsweringRequest,
	opts ...Option,
) ([]QuestionAnswering, error)

AnswerQuestion sends a question answering request and returns the answers.

The request must include both a question and a context. The model will identify the answer to the question within the provided context.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) AnswerTableQuestion

func (c Client) AnswerTableQuestion(
	req TableQuestionAnsweringRequest,
	opts ...Option,
) (TableQuestionAnswer, error)

AnswerTableQuestion sends a table question answering request and returns the answer.

The request must include both a question and a table. The model will identify the answer to the question within the provided table data.

NOTE: The HuggingFace API returns a bare JSON object for table question answering, not an array — despite the upstream schema declaring an array response. This method returns a single TableQuestionAnswer to match the actual API behavior.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) Chat

func (c Client) Chat(req ChatRequest, opts ...Option) (ChatResponse, error)

Chat sends a chat completion request and returns a chat completion response.

The request is passed by value and the SDK never mutates the received payload. The value copy shares the request's nested data (slices, maps, and pointed-to values) with the caller, so the caller must treat the request and the data it references as read-only while a call is in flight.

Concurrency:

  • A single Client is safe for concurrent use.
  • Reusing one request across sequential, fully-awaited calls is safe.
  • To invoke the same template request from multiple goroutines, pass a defensive copy per call, e.g. go client.Chat(req.Clone(), ...), or build a fresh request per call.

Model Precedence: The Model field is resolved with the following precedence (highest to lowest):

  1. ChatRequest.Model field (if non-nil and non-empty)
  2. Per-request options Model override
  3. Client-level Model option

Provider Precedence: The Provider field is applied as a fallback only if the resolved Model does not already contain a provider (indicated by ":" in the model string). If the Model is in the format "model:provider", the Provider option is ignored.

For example:

  • Model="mistral-7b", Provider="huggingface" → "mistral-7b:huggingface"
  • Model="mistral-7b:huggingface", Provider="mistral" → "mistral-7b:huggingface" (Provider ignored)
  • Model="mistral-7b:huggingface", Provider="" → "mistral-7b:huggingface"

Behavior:

  • Returns a configuration error if the request is missing a model or messages.
  • Returns a configuration error if *req.Stream is true; use ChatStream for streaming.

func (Client) ChatStream

func (c Client) ChatStream(req ChatRequest, opts ...Option) (*ChatStream, error)

ChatStream sends a chat completion request and returns a streaming response. Callers should Close the returned ChatStream when finished so the underlying HTTP connection and decoder goroutine are released promptly.

The request is passed by value and the SDK never mutates the received payload. The value copy shares the request's nested data (slices, maps, and pointed-to values) with the caller, so the caller must treat the request and the data it references as read-only while a call is in flight.

Concurrency:

  • A single Client is safe for concurrent use.
  • Reusing one request across sequential, fully-awaited calls is safe.
  • To invoke the same template request from multiple goroutines, pass a defensive copy per call, e.g. go client.ChatStream(req.Clone(), ...), or build a fresh request per call.

Model Precedence: The Model field is resolved with the following precedence (highest to lowest):

  1. ChatRequest.Model field (if non-nil and non-empty)
  2. Per-request options Model override
  3. Client-level Model option

Provider Precedence: The Provider field is applied as a fallback only if the resolved Model does not already contain a provider (indicated by ":" in the model string). If the Model is in the format "model:provider", the Provider option is ignored.

For example:

  • Model="mistral-7b", Provider="huggingface" → "mistral-7b:huggingface"
  • Model="mistral-7b:huggingface", Provider="mistral" → "mistral-7b:huggingface" (Provider ignored)
  • Model="mistral-7b:huggingface", Provider="" → "mistral-7b:huggingface"

Behavior:

  • Returns a configuration error if the request is missing a model or messages.
  • Always sends the request with streaming enabled.

func (Client) ClassifyText

func (c Client) ClassifyText(
	req TextClassificationRequest,
	opts ...Option,
) ([]TextClassification, error)

ClassifyText sends a text classification request and returns the text classification response for a single input.

For multiple classification inputs, use ClassifyTextBatch.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) ClassifyTextBatch

func (c Client) ClassifyTextBatch(
	req TextClassificationBatchRequest,
	opts ...Option,
) ([][]TextClassification, error)

ClassifyTextBatch sends a text classification request for a batch of inputs and returns a list of text classification responses for each input in the batch.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.

Callers should check the length of the response list before indexing.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) ClassifyTokens

func (c Client) ClassifyTokens(
	req TokenClassificationRequest,
	opts ...Option,
) ([]TokenClassification, error)

ClassifyTokens sends a token classification request and returns the token classification response for a single input.

For multiple inputs, use ClassifyTokensBatch.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) ClassifyTokensBatch

func (c Client) ClassifyTokensBatch(
	req TokenClassificationBatchRequest,
	opts ...Option,
) ([][]TokenClassification, error)

ClassifyTokensBatch sends a token classification request for a batch of inputs and returns a list of token classification responses for each input in the batch.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.

Callers should check the length of the response list before indexing.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) FillMask

func (c Client) FillMask(req FillMaskRequest, opts ...Option) ([]FillMaskPrediction, error)

FillMask sends a fill mask request and returns the mask filling predictions for a single input.

For multiple inputs, use FillMaskBatch.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) FillMaskBatch

func (c Client) FillMaskBatch(
	req FillMaskBatchRequest,
	opts ...Option,
) ([][]FillMaskPrediction, error)

FillMaskBatch sends a fill mask request for a batch of inputs and returns a list of mask filling predictions for each input in the batch.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.

Callers should check the length of the response list before indexing.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) Raw

func (c Client) Raw() RawService

Raw returns the raw HTTP request service for this client. Unlike the other endpoints, which are exposed directly as Client methods, the raw path remains namespaced under RawService: it is the advanced escape hatch for endpoints the SDK does not otherwise cover, and its several method variants are easier to discover grouped together than splashed across the Client surface.

RawService is immutable and captures a snapshot of the client options when created; it is lightweight, so prefer calling Raw() per use rather than retaining the value.

func (Client) Summarize

func (c Client) Summarize(req SummarizationRequest, opts ...Option) ([]Summarization, error)

Summarize sends a summarization request and returns the summarization output for a single input.

The API always returns a list for summarization; a single input yields a one-element list rather than a bare summary object.

For multiple inputs, use SummarizeBatch.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) SummarizeBatch

func (c Client) SummarizeBatch(
	req SummarizationBatchRequest,
	opts ...Option,
) ([]Summarization, error)

SummarizeBatch sends a summarization request for a batch of inputs and returns a flat list of summarization outputs, one for each input in the batch, in the same order as the inputs.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice. The response is a flat list (one summary per input) — not a nested list — consistent with how the API returns a list even for a single input.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) Translate

func (c Client) Translate(req TranslationRequest, opts ...Option) ([]Translation, error)

Translate sends a translation request and returns the translation output for a single input.

The API always returns a list for translation; a single input yields a one-element list rather than a bare translation object.

For multiple inputs, use TranslateBatch.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) TranslateBatch

func (c Client) TranslateBatch(
	req TranslationBatchRequest,
	opts ...Option,
) ([]Translation, error)

TranslateBatch sends a translation request for a batch of inputs and returns a flat list of translation outputs, one for each input in the batch, in the same order as the inputs.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice. The response is a flat list (one translation per input) — not a nested list — consistent with how the API returns a list even for a single input.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) ZeroShotClassifyText

func (c Client) ZeroShotClassifyText(
	req ZeroShotTextClassificationRequest,
	opts ...Option,
) ([]ZeroShotTextClassification, error)

ZeroShotClassifyText sends a zero-shot text classification request and returns the zero-shot text classification response for a single input.

For multiple inputs, use ZeroShotClassifyTextBatch.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

func (Client) ZeroShotClassifyTextBatch

func (c Client) ZeroShotClassifyTextBatch(
	req ZeroShotTextClassificationBatchRequest,
	opts ...Option,
) ([][]ZeroShotTextClassification, error)

ZeroShotClassifyTextBatch sends a zero-shot text classification request for a batch of inputs and returns a list of zero-shot text classification responses for each input in the batch.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.

Callers should check the length of the response list before indexing.

The Provider option is ignored for now, as hf-inference is currently the only supported provider.

type FillMaskBatchRequest

type FillMaskBatchRequest struct {
	// The inputs with masked tokens.
	// Required.
	Inputs []string `json:"inputs"`

	// Additional inference parameters for mask filling
	Parameters *FillMaskParameters `json:"parameters,omitempty"`
}

FillMaskBatchRequest represents a batched fill mask request to the API for multiple masked inputs.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.

func (*FillMaskBatchRequest) Clone

Clone returns a deep defensive copy of the request.

type FillMaskParameters

type FillMaskParameters struct {
	// When passed, the model will limit the scores to the passed targets instead of looking up
	// in the whole vocabulary. If the provided targets are not in the model vocab, they will be
	// tokenized and the first resulting token will be used (with a warning, and that might be
	// slower).
	Targets []string `json:"targets,omitempty"`

	// When passed, overrides the number of predictions to return.
	TopK *int `json:"top_k,omitempty"`
}

FillMaskParameters specify additional inference parameters for mask filling tasks.

func (*FillMaskParameters) Clone

Clone returns a deep defensive copy of the parameters.

type FillMaskPrediction

type FillMaskPrediction struct {
	// The input filled with the mask token prediction
	Sequence string `json:"sequence"`

	// The probability of the token prediction
	Score float64 `json:"score"`

	// The predicted token id (to replace the masked one).
	Token int `json:"token"`

	// The predicted token (to replace the masked one).
	TokenStr *string `json:"token_str"`
}

FillMaskPrediction represents a mask-filling output.

type FillMaskRequest

type FillMaskRequest struct {
	// The input text with masked tokens.
	// Required.
	Input string `json:"inputs"`

	// Additional inference parameters for mask filling
	Parameters *FillMaskParameters `json:"parameters,omitempty"`
}

FillMaskRequest represents a fill mask inference request to the API for a single input.

func (*FillMaskRequest) Clone

func (r *FillMaskRequest) Clone() FillMaskRequest

Clone returns a deep defensive copy of the request.

type MessageChunkType

type MessageChunkType string

MessageChunkType enumerates supported chat message chunk types.

const (
	// MessageChunkTypeText represents a text chunk.
	MessageChunkTypeText MessageChunkType = "text"
	// MessageChunkTypeImageURL represents an image_url chunk.
	MessageChunkTypeImageURL MessageChunkType = "image_url"
)

type Option

type Option = request.Option

Option represents a functional option that configures client requests.

func WithBaseURL

func WithBaseURL(u string) Option

WithBaseURL returns an Option that sets the base URL for API requests. The base URL is the root endpoint for all HuggingFace API calls and must not include query parameters or fragments.

func WithContext

func WithContext(ctx context.Context) Option

WithContext returns an Option that sets the context for API requests. The context can be used for cancellation, timeouts, and passing request-scoped values. If a nil context is provided, the SDK will fall back to context.Background().

func WithDefaultHTTPClient

func WithDefaultHTTPClient() Option

WithDefaultHTTPClient returns an Option that sets the default HTTP client.

func WithDefaultHeader

func WithDefaultHeader(key, value string) Option

WithDefaultHeader returns an Option that sets a header only if missing or empty.

func WithHTTPClientFactory

func WithHTTPClientFactory(factory func() http.Client) Option

WithHTTPClientFactory returns an Option that sets a http.Client created by the factory. The factory is invoked when request options are applied, so it can be used per request or at client construction time. The factory should return a fresh client value; avoid sharing mutable internals like Transport unless synchronized. If the factory is nil, the HTTP client is set to nil.

func WithHeader

func WithHeader(key, value string) Option

WithHeader returns an Option that sets a single header applied to every request.

func WithHeaders

func WithHeaders(h http.Header) Option

WithHeaders returns an Option that sets custom headers applied to every request, overriding any existing values for matching keys. Per-request headers can still override these values when provided.

func WithMaxResponseBodyBytes

func WithMaxResponseBodyBytes(n int64) Option

WithMaxResponseBodyBytes returns an Option that sets the maximum number of bytes read from any response body. Values <= 0 fall back to the default.

func WithModel

func WithModel(m string) Option

WithModel returns an Option that sets the model to use for API requests. The model specifies which HuggingFace model should process the request.

func WithProvider

func WithProvider(p string) Option

WithProvider returns an Option that sets the provider for API requests. The provider specifies which inference provider should handle the request.

func WithToken

func WithToken(t string) Option

WithToken returns an Option that sets the authentication token for API requests. The token is used for Bearer authentication with the HuggingFace API.

func WithUserAgentSuffix

func WithUserAgentSuffix(s string) Option

WithUserAgentSuffix returns an Option that appends a suffix to the SDK user agent string.

type QuestionAnswering

type QuestionAnswering struct {
	// The answer to the question.
	Answer string `json:"answer"`

	// The probability associated to the answer.
	Score float64 `json:"score"`

	// The character position in the input where the answer begins.
	Start int `json:"start"`

	// The character position in the input where the answer ends.
	End int `json:"end"`
}

QuestionAnswering represents a question answering output.

type QuestionAnsweringInput

type QuestionAnsweringInput struct {
	// The question to be answered.
	// Required.
	Question string `json:"question"`

	// The context to be used for answering the question.
	// Required.
	Context string `json:"context"`
}

QuestionAnsweringInput represents the input data for a question answering request. Both Question and Context are required.

type QuestionAnsweringParameters

type QuestionAnsweringParameters struct {
	// The number of answers to return (will be chosen by order of likelihood).
	// Note that less than top_k answers may be returned if there are not enough
	// options available within the context.
	TopK *int `json:"top_k,omitempty"`

	// If the context is too long to fit with the question for the model, it will
	// be split in several chunks with some overlap. This argument controls the
	// size of that overlap.
	DocStride *int `json:"doc_stride,omitempty"`

	// The maximum length of predicted answers (e.g., only answers with a shorter
	// length are considered).
	MaxAnswerLen *int `json:"max_answer_len,omitempty"`

	// The maximum length of the total sentence (context + question) in tokens of
	// each chunk passed to the model. The context will be split in several chunks
	// (using doc_stride as overlap) if needed.
	MaxSeqLen *int `json:"max_seq_len,omitempty"`

	// The maximum length of the question after tokenization. It will be truncated
	// if needed.
	MaxQuestionLen *int `json:"max_question_len,omitempty"`

	// Whether to accept impossible as an answer.
	HandleImpossibleAnswer *bool `json:"handle_impossible_answer,omitempty"`

	// Attempts to align the answer to real words. Improves quality on space
	// separated languages. Might hurt on non-space-separated languages (like
	// Japanese or Chinese).
	AlignToWords *bool `json:"align_to_words,omitempty"`
}

QuestionAnsweringParameters specify additional inference parameters for question answering.

func (*QuestionAnsweringParameters) Clone

Clone returns a deep defensive copy of the parameters.

type QuestionAnsweringRequest

type QuestionAnsweringRequest struct {
	// The question and context pair to answer.
	// Required.
	Input QuestionAnsweringInput `json:"inputs"`

	// Additional inference parameters for question answering.
	Parameters *QuestionAnsweringParameters `json:"parameters,omitempty"`
}

QuestionAnsweringRequest represents a question answering inference request to the API.

func (*QuestionAnsweringRequest) Clone

Clone returns a deep defensive copy of the request.

type RawEvent

type RawEvent struct {
	Data  []byte
	Event string
	ID    string
	Retry *time.Duration
}

RawEvent mirrors the SSE fields returned by raw streams.

type RawService

type RawService struct {
	// contains filtered or unexported fields
}

RawService sends raw HTTP requests using the configured request options.

It is the deliberate exception to the rest of the SDK, where endpoints are exposed as Client methods: RawService is the advanced escape hatch for endpoints the SDK does not model type-safely. It combines several axes (byte-slice or io.Reader bodies, typed error handling or raw HTTP responses, one-shot or SSE streaming) into eight methods, which is enough surface that keeping it namespaced under Client.Raw() avoids cluttering the Client API.

RawService is immutable and holds a snapshot of the client options.

func (RawService) Do

func (r RawService) Do(
	requestBody []byte,
	method string,
	path string,
	opts ...Option,
) (*http.Response, error)

Do performs a raw HTTP request with a byte slice body and applies SDK error interpretation on non-2xx responses. The caller must close resp.Body on success.

func (RawService) DoRaw

func (r RawService) DoRaw(
	requestBody []byte,
	method string,
	path string,
	opts ...Option,
) (*http.Response, error)

DoRaw performs a raw HTTP request with a byte slice body without translating non-2xx responses into SDK errors. The caller must close resp.Body on success.

func (RawService) DoRawReader

func (r RawService) DoRawReader(
	requestBody io.Reader,
	method string,
	path string,
	opts ...Option,
) (*http.Response, error)

DoRawReader performs a raw HTTP request with a streaming body without translating non-2xx responses into SDK errors. The caller must close resp.Body on success.

func (RawService) DoReader

func (r RawService) DoReader(
	requestBody io.Reader,
	method string,
	path string,
	opts ...Option,
) (*http.Response, error)

DoReader performs a raw HTTP request with a streaming body and applies SDK error interpretation on non-2xx responses. The caller must close resp.Body on success.

func (RawService) Stream

func (r RawService) Stream(
	requestBody []byte,
	method string,
	path string,
	opts ...Option,
) (*RawStream, error)

Stream performs a raw HTTP request and returns an SSE stream, applying SDK error interpretation on non-2xx responses. Callers should Close the returned RawStream when finished to promptly release the HTTP connection and decoder goroutine.

func (RawService) StreamRaw

func (r RawService) StreamRaw(
	requestBody []byte,
	method string,
	path string,
	opts ...Option,
) (*RawStream, error)

StreamRaw performs a raw HTTP request and returns an SSE stream without translating non-2xx responses into SDK errors. This function is probably only interesting to advanced users. Only use this when you need to inspect the raw response; callers are responsible for interpreting HTTP errors themselves. Callers should Close the returned RawStream when finished to promptly release the HTTP connection and decoder goroutine.

func (RawService) StreamRawReader

func (r RawService) StreamRawReader(
	requestBody io.Reader,
	method string,
	path string,
	opts ...Option,
) (*RawStream, error)

StreamRawReader performs a raw HTTP request with a streaming body and returns an SSE stream without translating non-2xx responses into SDK errors. This function is probably only interesting to advanced users. Only use this when you need to inspect the raw response; callers are responsible for interpreting HTTP errors themselves. Callers should Close the returned RawStream when finished to promptly release the HTTP connection and decoder goroutine.

func (RawService) StreamReader

func (r RawService) StreamReader(
	requestBody io.Reader,
	method string,
	path string,
	opts ...Option,
) (*RawStream, error)

StreamReader performs a raw HTTP request with a streaming body and returns an SSE stream with SDK error interpretation. Callers should Close the returned RawStream when finished to promptly release the HTTP connection and decoder goroutine.

type RawStream

type RawStream struct {
	// contains filtered or unexported fields
}

RawStream exposes a raw SSE stream returned by RawService stream methods.

func (*RawStream) Close

func (s *RawStream) Close() error

Close releases the underlying stream resources.

func (*RawStream) Recv

func (s *RawStream) Recv(ctx context.Context) (event RawEvent, err error)

Recv blocks until the next SSE event is available or the context is done.

type ResponseFormatType

type ResponseFormatType string

ResponseFormatType enumerates known response formats.

const (
	// ResponseFormatTypeText requests a text response.
	ResponseFormatTypeText ResponseFormatType = "text"
	// ResponseFormatTypeJSONSchema requests a JSON schema response.
	ResponseFormatTypeJSONSchema ResponseFormatType = "json_schema"
	// ResponseFormatTypeJSONObject requests a JSON object response.
	ResponseFormatTypeJSONObject ResponseFormatType = "json_object"
)

type SDKError

type SDKError = hferrors.SDKError

SDKError represents a client-side SDK error that occurred before a response was received from the API.

Users can type-assert errors to *SDKError to access the error kind and underlying cause:

if sdkErr, ok := err.(*hfgo.SDKError); ok {
    fmt.Printf("Kind %s: %s\n", sdkErr.Kind, sdkErr.Message)
}

type SDKErrorKind

type SDKErrorKind = hferrors.SDKErrorKind

SDKErrorKind represents the category of a client-side SDK error.

type Summarization

type Summarization struct {
	// The summarized text.
	SummaryText string `json:"summary_text"`
}

Summarization represents a summarization output.

type SummarizationBatchRequest

type SummarizationBatchRequest struct {
	// The texts to summarize.
	// Required.
	Inputs []string `json:"inputs"`

	// Additional inference parameters for summarization.
	Parameters *SummarizationParameters `json:"parameters,omitempty"`
}

SummarizationBatchRequest represents a batched summarization inference request to the API for multiple inputs.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.

func (*SummarizationBatchRequest) Clone

Clone returns a deep defensive copy of the request.

type SummarizationParameters

type SummarizationParameters struct {
	// Whether to clean up the potential extra spaces in the text output.
	CleanUpTokenizationSpaces *bool `json:"clean_up_tokenization_spaces,omitempty"`

	// The truncation strategy to use.
	Truncation *string `json:"truncation,omitempty"`

	// GenerateParameters provides additional parametrization of the text
	// generation algorithm (e.g. "max_new_tokens", "temperature", "top_k").
	//
	// It is exposed because it is part of the upstream inference schema, but the
	// set of valid arguments is not documented by Hugging Face and may depend on
	// the model being used. Invalid or unsupported arguments can be rejected by
	// the API.
	GenerateParameters map[string]any `json:"generate_parameters,omitempty"`
}

SummarizationParameters specify additional inference parameters for summarization tasks.

func (*SummarizationParameters) Clone

Clone returns a deep defensive copy of the parameters. The GenerateParameters map is copied as a new map, but its values are shared because their types are not known statically.

type SummarizationRequest

type SummarizationRequest struct {
	// The text to summarize.
	// Required.
	Input string `json:"inputs"`

	// Additional inference parameters for summarization.
	Parameters *SummarizationParameters `json:"parameters,omitempty"`
}

SummarizationRequest represents a summarization inference request to the API for a single input.

func (*SummarizationRequest) Clone

Clone returns a deep defensive copy of the request.

type TableQuestionAnswer

type TableQuestionAnswer struct {
	// The answer to the question given the table. If there is an aggregator,
	// the answer will be preceded by "AGGREGATOR >".
	Answer string `json:"answer"`

	// Coordinates of the cells of the answers.
	Coordinates [][]int `json:"coordinates"`

	// List of strings made up of the answer cell values.
	Cells []string `json:"cells"`

	// If the model has an aggregator, this returns the aggregator.
	Aggregator *string `json:"aggregator,omitempty"`
}

TableQuestionAnswer represents a table question answering output.

type TableQuestionAnsweringInput

type TableQuestionAnsweringInput struct {
	// The question to be answered about the table.
	// Required.
	Question string `json:"question"`

	// The table to serve as context for the questions.
	// Each key is a column name, and the value is a list of cell values for that column.
	// Required.
	Table map[string][]string `json:"table"`
}

TableQuestionAnsweringInput represents the input data for a table question answering request. Both Question and Table are required.

type TableQuestionAnsweringParameters

type TableQuestionAnsweringParameters struct {
	// Activates and controls padding.
	Padding *string `json:"padding,omitempty"`

	// Whether to do inference sequentially or as a batch. Batching is faster,
	// but models like SQA require the inference to be done sequentially to
	// extract relations within sequences, given their conversational nature.
	Sequential *bool `json:"sequential,omitempty"`

	// Activates and controls truncation.
	Truncation *bool `json:"truncation,omitempty"`
}

TableQuestionAnsweringParameters specify additional inference parameters for table question answering tasks.

func (*TableQuestionAnsweringParameters) Clone

Clone returns a deep defensive copy of the parameters.

type TableQuestionAnsweringRequest

type TableQuestionAnsweringRequest struct {
	// The question and table pair to answer.
	// Required.
	Input TableQuestionAnsweringInput `json:"inputs"`

	// Additional inference parameters for table question answering.
	Parameters *TableQuestionAnsweringParameters `json:"parameters,omitempty"`
}

TableQuestionAnsweringRequest represents a table question answering inference request to the API.

func (*TableQuestionAnsweringRequest) Clone

Clone returns a deep defensive copy of the request.

type TextClassification

type TextClassification struct {
	// The predicted class label.
	Label string `json:"label"`

	// The corresponding probability.
	Score float64 `json:"score"`
}

TextClassification represents a text classification output.

type TextClassificationBatchRequest

type TextClassificationBatchRequest struct {
	// The texts to classify.
	// Required.
	Inputs []string `json:"inputs"`

	// Additional inference parameters for text classification
	Parameters *TextClassificationParameters `json:"parameters,omitempty"`
}

TextClassificationBatchRequest represents a batched text classification inference request to the API for multiple inputs.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.

func (*TextClassificationBatchRequest) Clone

Clone returns a deep defensive copy of the request.

type TextClassificationParameters

type TextClassificationParameters struct {
	// Possible values: sigmoid, softmax, none.
	FunctionToApply *string `json:"function_to_apply,omitempty"`

	// When specified, limits the output to the top K most probable classes.
	TopK *int `json:"top_k,omitempty"`
}

TextClassificationParameters specify additional inference parameters for text classification.

func (*TextClassificationParameters) Clone

Clone returns a deep defensive copy of the parameters.

type TextClassificationRequest

type TextClassificationRequest struct {
	// The text to classify.
	// Required.
	Input string `json:"inputs"`

	// Additional inference parameters for text classification
	Parameters *TextClassificationParameters `json:"parameters,omitempty"`
}

TextClassificationRequest represents a text classification inference request to the API for a single input.

func (*TextClassificationRequest) Clone

Clone returns a deep defensive copy of the request.

type TokenClassification

type TokenClassification struct {
	// The predicted label for a group of one or more tokens.
	// Present when an aggregation strategy other than "none" is used.
	EntityGroup *string `json:"entity_group,omitempty"`

	// The predicted label for a single token.
	// Present when aggregation_strategy is "none".
	Entity *string `json:"entity,omitempty"`

	// The associated score / probability.
	Score float64 `json:"score"`

	// The corresponding text.
	Word string `json:"word"`

	// The character position in the input where this group begins.
	Start int `json:"start"`

	// The character position in the input where this group ends.
	End int `json:"end"`
}

TokenClassification represents a token classification output.

type TokenClassificationBatchRequest

type TokenClassificationBatchRequest struct {
	// The texts to classify tokens from.
	// Required.
	Inputs []string `json:"inputs"`

	// Additional inference parameters for token classification.
	Parameters *TokenClassificationParameters `json:"parameters,omitempty"`
}

TokenClassificationBatchRequest represents a batched token classification request to the API for multiple inputs.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.

func (*TokenClassificationBatchRequest) Clone

Clone returns a deep defensive copy of the request.

type TokenClassificationParameters

type TokenClassificationParameters struct {
	// A list of labels to ignore in the classification results.
	IgnoreLabels []string `json:"ignore_labels,omitempty"`

	// The number of overlapping tokens between chunks when splitting the input text.
	Stride *int `json:"stride,omitempty"`

	// The strategy used to fuse tokens based on model predictions.
	// Possible values: "none", "simple", "first", "average", "max".
	AggregationStrategy *string `json:"aggregation_strategy,omitempty"`
}

TokenClassificationParameters specify additional inference parameters for token classification.

func (*TokenClassificationParameters) Clone

Clone returns a deep defensive copy of the parameters.

type TokenClassificationRequest

type TokenClassificationRequest struct {
	// The text to classify tokens from.
	// Required.
	Input string `json:"inputs"`

	// Additional inference parameters for token classification.
	Parameters *TokenClassificationParameters `json:"parameters,omitempty"`
}

TokenClassificationRequest represents a token classification inference request to the API for a single input.

func (*TokenClassificationRequest) Clone

Clone returns a deep defensive copy of the request.

type ToolChoiceMode

type ToolChoiceMode string

ToolChoiceMode enumerates known tool choice modes.

const (
	// ToolChoiceModeAuto lets the provider decide the tool choice.
	ToolChoiceModeAuto ToolChoiceMode = "auto"
	// ToolChoiceModeNone disables tool usage.
	ToolChoiceModeNone ToolChoiceMode = "none"
	// ToolChoiceModeRequired requires tool usage.
	ToolChoiceModeRequired ToolChoiceMode = "required"
)

type Translation

type Translation struct {
	// The translated text.
	TranslationText string `json:"translation_text"`
}

Translation represents a translation output.

type TranslationBatchRequest

type TranslationBatchRequest struct {
	// The texts to translate.
	// Required.
	Inputs []string `json:"inputs"`

	// Additional inference parameters for translation.
	Parameters *TranslationParameters `json:"parameters,omitempty"`
}

TranslationBatchRequest represents a batched translation inference request to the API for multiple inputs.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.

func (*TranslationBatchRequest) Clone

Clone returns a deep defensive copy of the request.

type TranslationParameters

type TranslationParameters struct {
	// Whether to clean up the potential extra spaces in the text output.
	CleanUpTokenizationSpaces *bool `json:"clean_up_tokenization_spaces,omitempty"`

	// The source language of the text. Required for models that can
	// translate from multiple languages.
	SrcLang *string `json:"src_lang,omitempty"`

	// The target language to translate to. Required for models that can
	// translate to multiple languages.
	TgtLang *string `json:"tgt_lang,omitempty"`

	// The truncation strategy to use.
	Truncation *string `json:"truncation,omitempty"`

	// GenerateParameters provides additional parametrization of the text
	// generation algorithm (e.g. "max_new_tokens", "temperature", "top_k").
	//
	// It is exposed because it is part of the upstream inference schema, but the
	// set of valid arguments is not documented by Hugging Face and may depend on
	// the model being used. Invalid or unsupported arguments can be rejected by
	// the API.
	GenerateParameters map[string]any `json:"generate_parameters,omitempty"`
}

TranslationParameters specify additional inference parameters for translation tasks.

func (*TranslationParameters) Clone

Clone returns a deep defensive copy of the parameters. The GenerateParameters map is copied as a new map, but its values are shared because their types are not known statically.

type TranslationRequest

type TranslationRequest struct {
	// The text to translate.
	// Required.
	Input string `json:"inputs"`

	// Additional inference parameters for translation.
	Parameters *TranslationParameters `json:"parameters,omitempty"`
}

TranslationRequest represents a translation inference request to the API for a single input.

func (*TranslationRequest) Clone

Clone returns a deep defensive copy of the request.

type ZeroShotTextClassification

type ZeroShotTextClassification struct {
	// The predicted class label.
	Label string `json:"label"`

	// The corresponding probability.
	Score float64 `json:"score"`
}

ZeroShotTextClassification represents a zero-shot text classification output.

type ZeroShotTextClassificationBatchRequest

type ZeroShotTextClassificationBatchRequest struct {
	// The texts to classify.
	// Required.
	Inputs []string `json:"inputs"`

	// Additional inference parameters for zero-shot text classification.
	// Required.
	Parameters *ZeroShotTextClassificationParameters `json:"parameters,omitempty"`
}

ZeroShotTextClassificationBatchRequest represents a batched zero-shot text classification request to the API for multiple inputs.

NOTE: Batched inference is supported by the upstream API, but is not officially documented; behavior may change without notice.

func (*ZeroShotTextClassificationBatchRequest) Clone

Clone returns a deep defensive copy of the request.

type ZeroShotTextClassificationParameters

type ZeroShotTextClassificationParameters struct {
	// The set of possible class labels to classify the text into.
	// Required.
	CandidateLabels []string `json:"candidate_labels,omitempty"`

	// The sentence used in conjunction with candidate_labels to attempt
	// the text classification by replacing the placeholder with the
	// candidate labels.
	HypothesisTemplate *string `json:"hypothesis_template,omitempty"`

	// Whether multiple candidate labels can be true. If false, the scores
	// are normalized such that the sum of the label likelihoods for each
	// sequence is 1. If true, the labels are considered independent and
	// probabilities are normalized for each candidate.
	MultiLabel *bool `json:"multi_label,omitempty"`
}

ZeroShotTextClassificationParameters specify additional inference parameters for zero-shot text classification.

func (*ZeroShotTextClassificationParameters) Clone

Clone returns a deep defensive copy of the parameters.

type ZeroShotTextClassificationRequest

type ZeroShotTextClassificationRequest struct {
	// The text to classify.
	// Required.
	Input string `json:"inputs"`

	// Additional inference parameters for zero-shot text classification.
	// Required.
	Parameters *ZeroShotTextClassificationParameters `json:"parameters,omitempty"`
}

ZeroShotTextClassificationRequest represents a zero-shot text classification request to the API for a single input.

func (*ZeroShotTextClassificationRequest) Clone

Clone returns a deep defensive copy of the request.

Directories

Path Synopsis
examples
chat/basic command
chat/convo command
chat/streaming command
fill-mask/basic command
fill-mask/batch command
internal
chatstream
Package chatstream provides helpers for working with streamed chat responses.
Package chatstream provides helpers for working with streamed chat responses.
hferrors
Package hferrors defines the reusable SDK error types returned by hfgo.
Package hferrors defines the reusable SDK error types returned by hfgo.
request
Package request contains the lower-level HTTP, SSE, and JSON utilities used by the SDK to build, send, and process Hugging Face API requests.
Package request contains the lower-level HTTP, SSE, and JSON utilities used by the SDK to build, send, and process Hugging Face API requests.
sdkversion
Package sdkversion exposes the SDK version string used in User-Agent headers.
Package sdkversion exposes the SDK version string used in User-Agent headers.
testutils
Package testutils provides helpers for tests.
Package testutils provides helpers for tests.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL