* test(relayconvert): add golden snapshot matrix and relaykit boundary guard
Phase 0 of the relaykit extraction plan: pin byte-level output of every
registered (from,to) request/response/stream conversion route, and
forbid kit-bound packages from growing host-only imports.
* wip(relayconvert): drop gin.Context from converter signatures; add convmeta draft
Phase 1 in progress: relayconvert now takes context.Context; host media
resolver adapts gin.Context back at the service boundary.
* refactor(relayconvert): decouple converters from RelayInfo, gin, and settings
Phase 1 of the relaykit extraction plan:
- converters now depend on convmeta.Meta (implemented by RelayInfo) instead
of *relaycommon.RelayInfo; ClaudeConvertInfo and the format guesser move
to convmeta with aliases left behind
- host settings reach converters via a convmeta.Options snapshot built in
RelayInfo.ConvOptions; no more model_setting/reasoning global reads inside
the conversion layer
- effort-suffix helpers move to service/relayconvert/reasoning (old package
forwards); chat-to-responses upgrade policy moves to service (host routing
logic, not conversion)
- golden conversion matrix unchanged
* test(relayconvert): tighten boundary — kit packages now free of gin/setting imports
* refactor(dto): drop gin and logger dependencies
Phase 2 (part 1): dto.Request.IsStream now takes *http.Request instead of
*gin.Context (Gemini's impl reads query/path off the std request); dto's
three logger calls become common.SysError. Boundary test allowlist is now
empty — kit-bound packages import no gin/setting/logger/model.
* refactor(kit): extract dependency-free kitutil; dto/types/relayconvert stop importing common
Phase 2 of the relaykit extraction plan:
- new service/relayconvert/kitutil holds the pure helpers the kit needs
(JSON wrappers, pointer/string/uuid/timestamp utils, MaskSensitiveInfo,
pluggable LogInfo/LogError hooks, Debug flag)
- dto, types, and all relayconvert packages now use kitutil; their only
remaining internal deps are dto/types/constant
- common keeps every original symbol (MaskSensitiveInfo delegates to
kitutil) so host code is untouched; main.go routes kit logging into
common.SysLog/SysError and mirrors DebugEnabled
- golden conversion matrix unchanged
* refactor(kit): move EndpointType/FinishReason to types; OpenRouter dialect via Options
Kit packages (dto/types/relayconvert/reasonmap) no longer import constant:
- EndpointType and finish-reason values live in types; constant re-exports
- the OpenRouter special-case in claude->openai request conversion reads
Options.OpenRouterDialect, set by the host from the channel type;
InitChannelMeta invalidates the cached snapshot on channel switch
* refactor: extract relaykit submodule (dto/types/relayconvert/reasonmap)
Phase 3 of the relaykit extraction plan:
- new go module github.com/QuantumNous/new-api/relaykit containing dto
(minus task family), types, relayconvert (with convmeta/kitutil/reasoning),
and reasonmap; host consumes it via require + replace, go.work for dev
- task-family dto (task/suno/midjourney/video) stays in the host dto
package; dual-consumer host files alias it as taskdto
- relaykit builds and tests standalone (GOWORK=off): no host imports,
no gin, no DB, no settings
- golden conversion matrix unchanged
* build(docker): copy relaykit/go.mod before go mod download
The local-replace submodule's go.mod must exist inside the build context
for the main module graph to resolve.
* fix: address relaykit extraction regressions
* fix: address relaykit review regressions
* docs: document Meta nil receiver contract
* fix(relaykit): fail OpenAI→Claude conversion without max_tokens; reject negative default_max_tokens
The Claude Messages API requires max_tokens (omitting it is a 400
"Field required"), but with a nil Options.Claude.DefaultMaxTokens hook
the converters silently emitted a request the upstream is guaranteed to
reject. Both OpenAI Chat and Responses → Claude conversions now return
sharedclaude.ErrMissingMaxTokens when no path (client value, default
hook, thinking-adapter floor) supplied one. Unreachable in the host,
which always configures the hook.
Host side, claude.default_max_tokens now rejects negative values at the
option API before persisting — they would wrap into huge unsigned values
during conversion. Zero stays allowed: the current API treats
max_tokens: 0 as cache pre-warming.
* fix: make Gemini safety settings read path race-free
The line-number layer spans the full editor, so giving it an opaque
background and raised z-index covered the highlighted content. Reverts
the gutter background added in 88868002f; the horizontal-scroll
gutter overlap it targeted returns as a known cosmetic issue.
When the dashboard token refresh endpoint returned 429 (shared
critical rate limit) the frontend classified it as out_of_sync,
cleared local auth state, and redirected to /sign-in. The rate limit
itself is working as intended; the bug is that a temporary rejection
was treated as a terminal auth failure.
- Treat 429 refresh responses as transient errors on the frontend,
keeping the session retryable instead of clearing it. Only explicit
401 or confirmed session mismatch/race exhaustion clears auth state.
- Return Retry-After on all rate-limited responses (remaining TTL on
Redis, window duration on the in-memory limiter) so clients can
back off.
- Log the underlying error with request context when auth session
errors map to 500 AUTH_INTERNAL_ERROR, and replace fmt.Println with
request-scoped logging in the Redis rate limiter error paths.
Fixes#6361
Persist committed editor drafts to the source model during batch copy so
targets never carry pricing the source would lose, surface loading/error
states for the enabled-models query, and compare saved* props in the
visual editor memo so rows leave the unset list reliably after save.
Parse OpenAI's native cache_write_tokens (chat prompt_tokens_details /
responses input_tokens_details), bill it at the cache-creation ratio, and
clamp the uncached prompt remainder at zero since cached + cache-write can
exceed prompt_tokens. Propagate the field through chat/responses/claude
format conversions and tiered expression billing (cc variable).
* refactor: consolidate relay protocol converters
* refactor relayconvert text converters
* feat: refine relay converters and advanced custom routing
* refactor: enhance logging and add thought signature handling for Gemini requests
* refactor: enhance channel cache and pricing endpoint handling for advanced custom models
* feat: preserve billing usage semantics
* feat: add protocol-aware billing usage
* Delete useless files
* chore: update action versions in workflow files
* chore: update Docker action versions in workflow files
* fix: harden billing usage settlement and hot-path route matching
- estimate Gemini completion tokens locally when billable usageMetadata is
prompt-only but output content was received (e.g. client aborts the stream
before the final chunk), and rebuild the attached billing_usage as estimated
so settlement does not bill zero output tokens
- guard NewClaudeMessagesBillingUsage against all-zero ClaudeUsage, matching
the OpenAI/Gemini constructors, so a zero billing_usage cannot override a
non-zero top-level usage during settlement
- cache compiled advanced-custom route model regexes; they run on the request
hot path and were recompiled per request
- move the effectiveBillingUsage remap to PostTextConsumeQuota only, and
document that calculateTextQuotaSummary expects remapped usage
- document the updatePricingLock -> channelSyncLock lock ordering that
InitChannelCache/CacheUpdateChannel rely on, and the aux-struct pitfall in
GeminiChatResponse.UnmarshalJSON
Revert 6869cd94b (perf(web): align table badge spacing).
Restores px padding on status badges so usage-log timing/duration
pills are not flush against the edges.
Thread int32 saturation clamps from tiered settlement and video task
recompute into the consume/task logs under admin_info, so oversized or
malformed billing inputs stay auditable. Clamp negative audio duration
before token conversion and gate the saturation UI markers on admin.
Bound max-tokens fields across all relay format validators, saturate
tiered-expression rounding and audio/tool/task token conversions, and
route legacy remix ratios through the guarded setter.
Bound user-supplied count/duration parameters at request validation,
route ratio multipliers through guarded setters, and use saturating
int conversions in all quota math paths.
Pin down the checkUpdatePassword contract:
- changing a password requires the correct current password
- OAuth/passwordless accounts (empty password hash) cannot set a
password through the self-service endpoint and must use the password
reset flow instead
- setupLogin never writes back the password column
These guard against regressing the password-change path back into a
short-circuit that lets passwordless accounts set a password directly.
- add SESSION_COOKIE_SECURE / SESSION_COOKIE_TRUSTED_URL env vars with
startup validation: enabling Secure requires at least one trusted
HTTPS entry URL
- wire common.SessionCookieSecure into the session cookie store instead
of a hardcoded Secure=false
- print a startup warning when Secure session cookies are disabled
- document the new settings in .env.example and docker-compose files
Secure stays off by default because many deployments front new-api with
plain-HTTP reverse proxies, where a hardcoded Secure default would break
logins entirely; enabling it safely depends on the deployment's TLS
setup, so it ships as an opt-in deployment-hardening flag.
- normalize emails (trim + lowercase) and enforce uniqueness across
registration, OAuth auto-registration, and email binding
- serialize concurrent writers on the same normalized email within a
transaction to avoid duplicate accounts
- resolve password reset to a single matching account and reject
ambiguous or absent matches
- require an existing password before self-service password change and
reject login for accounts without a usable password
* fix(openai): harden Chat-to-Responses compatibility
Add a shared Responses-to-Chat stream state machine and use it from the OpenAI relay path. Preserve assistant text alongside tool calls, bind tool argument deltas by output_index, map incomplete finish reasons, support reasoning/custom tool events, and buffer upstream SSE for non-stream Chat clients.
Add deterministic service tests and relay SSE tests for the conversion path.
Related to #5745.
* refactor: rename openaicompat to relayconvert for improved clarity
* feat(gemini): support responses request conversion
* feat: add responses to chat conversion support
* fix: harden responses chat conversion edge cases
Add a shared Responses-to-Chat stream state machine and use it from the OpenAI relay path. Preserve assistant text alongside tool calls, bind tool argument deltas by output_index, map incomplete finish reasons, support reasoning/custom tool events, and buffer upstream SSE for non-stream Chat clients.
Add deterministic service tests and relay SSE tests for the conversion path.
Related to #5745.
* feat: add system instance reporting
* feat: show system instance resources
* fix: update translations for heartbeat messages in Russian and Vietnamese
- add vercel-react-best-practices skill (SKILL.md + full-guide.md)
- slim CLAUDE.md to import shared AGENTS.md conventions
- promote go-ntlmssp to a direct dependency in go.mod
Replace the per-batch ClickHouse mutation loop with a single
ALTER TABLE ... DELETE, since ClickHouse DELETE is a heavy mutation
that rewrites data parts and per-batch mutations are pathologically
slow. Add deterministic unit tests covering ClickHouse DSN handling,
main-database rejection, TTL DDL generation, log ordering, request_id
backfill, and display id assignment.
Add an eye toggle in the flow section header that masks sensitive node
labels (users, tokens, nodes, groups, channels) in the Sankey while keeping
model names visible. Masking only rewrites display text; nodes stay distinct
via their key so graph structure, links, and highlighting are unaffected.
- Highlight full paths through a clicked node or link in the flow Sankey,
dimming unrelated nodes/links instead of removing them
- Disable VChart built-in emphasis to avoid crash, use custom highlight sets
- Initialize models filter dialog from currently applied filters so manual
time ranges are not overridden by preferences; auto-pick granularity by range
- Lift user charts time range/granularity/limit to dashboard as controlled state
- Show 3-column card grid from xl breakpoint instead of 2xl
- Cap inline priority/weight width to avoid huge values stretching cards
- Collapse right-column grid to content-sized columns, removing wasted space
- Memoize channel columns, context value, upstream-update result, and ChannelCard
to avoid rebuilding/re-rendering all cards on unrelated state changes