Dropping the BFF: an admin SPA that talks to Spring Boot directly
Moving Saramin Vietnam's admin from Next.js and Prisma to a Vite + antd SPA on httpOnly cookies and CSRF, and three auth bugs that passed every typecheck.

On this page
The admin console for Saramin Vietnam started life as a Next.js app with its own backend-for-frontend. It had 72 route handlers, about 2,300 lines of them, and almost every one re-declared an endpoint that the Spring Boot backend already had. Login went through the BFF, which read the backend's Set-Cookie headers and set the same values again as its own cookies. Every list page went through a handler that forwarded the request and reshaped the answer.
When I wrote the migration plan, the question I kept coming back to was simple: what is this layer protecting? The answer turned out to be "an origin boundary". The admin and the API lived on different origins, so the browser could not hold the backend's cookies for the admin, and the BFF existed to paper over that. Put both behind one origin and the whole ceremony goes away without losing anything. The backend's own cookie code already said it expected the admin to be same-site.
So we dropped it. This post covers what replaced it, and the three auth bugs I found only by clicking through the result in a real browser, after tsc, lint and the build had all passed.
What there was to delete
The plan started by measuring the old app instead of guessing at it. It was about 49,700 lines across 360 files: Next.js 16, next-intl, Prisma over a SQLite prototype, shadcn and Base UI, react-hook-form. The BFF itself, meaning the route handlers plus the token, proxy and audit helpers, was only 6.2% of that.
That number reframed the project. Removing the BFF was the cheap part. The other 93.8% was UI, and moving to a plain SPA with antd meant rewriting it. The team chose antd over keeping shadcn, accepting a different component library than the rest of the company's apps used. Two things came across untouched: the message catalogs (2,630 keys in three locales) and the date-formatting helper that pins the Vietnam wall clock, because it already carried a fixed bug nobody wanted to fix twice.
The migration landed on 2026-08-01 as one commit: 781 files changed, about 23,000 lines added and 72,000 removed, with all nine tracked modules ported. Screens that only worked because Prisma had a local table behind them did not come across. They became feature-flagged modules with no API, waiting for the backend to grow the endpoint.
The new shape

In production, one nginx serves the built dist/ at / and reverse-proxies /api/* to the Spring Boot service. The browser sees one origin, so the backend's cookies are first-party for the admin:
access_token: httpOnly, SameSite=Lax, 30 minutes.refresh_token: httpOnly, SameSite=Lax, 14 days.csrf_token: readable by JavaScript, SameSite=Lax, 14 days.
The SPA never stores a token. There is no Authorization header, nothing in localStorage, nothing in a Zustand store. All calls go through one axios client with withCredentials: true, and that client is the only place that knows about auth.
CSRF is double-submit. For unsafe methods the client reads the csrf_token cookie and echoes it in a header:
api.interceptors.request.use((config) => {
const method = (config.method ?? 'get').toUpperCase();
if (UNSAFE.has(method)) {
const token = readCookie('csrf_token');
if (token) config.headers['X-CSRF-TOKEN'] = token; // raw value, not masked
}
return config;
});The "raw" part matters. Spring Security has a handler that expects an XOR-masked token and one that expects the plain value. The backend uses the plain one, so masking it on the client gets every write rejected.
The refresh flow is single-flight. When several requests fail with 401 at once, they all wait on one refresh call instead of each starting their own:
let refreshInFlight: Promise<void> | null = null;
async function refreshOnce() {
refreshInFlight ??= rawAxios.post('/api/auth/refresh').finally(() => {
refreshInFlight = null;
});
return refreshInFlight;
}Only the backend's "unauthorized" and "invalid token" error codes trigger a refresh. A wrong password and an inactive account also return 401, but retrying those would be pointless. Login, refresh, logout and set-password never refresh at all, each request retries once, and the refresh call itself uses bare axios so it can't re-enter the interceptor. A rejected CSRF token gets one refresh too, because the refresh mints a new CSRF cookie.
One nginx trap worth knowing
The first version of the nginx config proxied to the backend's service name directly. After a backend redeploy, every /api/* call returned 502. nginx resolves an upstream hostname once at startup and keeps that IP, and the old container's IP was gone. The fix is to put the upstream in a variable and give nginx a resolver with a short valid, which makes it resolve per request:
location /api/ {
resolver 127.0.0.11 valid=10s; # Docker's embedded DNS
set $upstream http://backend:8080;
proxy_pass $upstream$request_uri;
}The SPA side is the usual try_files $uri /index.html, with index.html served no-cache and hashed /assets/ served immutable.
Three bugs that passed every check
All three of these were caught while I was verifying the migration by hand, before the migration commit landed. None of them is visible to the type system. That's the reason the repo now has a rule that says so in its own words: typecheck passing is not verification.

1. The session-refresh reload loop
The first version of the refresh-failure handler did what felt natural: window.location.assign('/login'). On a logged-out cold visit that goes like this. /me returns 401, so the client tries a refresh, which also returns 401, so the page hard-reloads to /login. The app boots again, asks /me again, gets 401 again, and reloads again. Forever.
The fix is to never reload. When the refresh fails, the interceptor writes null into the current-user query and throws:
} catch {
queryClient.setQueryData(ME_QUERY_KEY, null);
throw new ApiError('UNAUTHORIZED');
}React Router then redirects from inside the running app, and the cycle has nowhere to restart from. When I later added handling for a "this token belongs to a different user" error, I reused exactly this path on purpose.
2. The logout redirect race
logout() called navigate('/login', { replace: true }). Meanwhile ProtectedRoute rendered <Navigate to="/login" replace state={{ from }} /> as soon as the user became null. Both reacted to the same state change, and whichever ran last won. Which one that was depended on timing, so logout behaved differently from one click to the next.
The rule that came out of it is now the second non-negotiable in the repo's agent instructions: one redirect mechanism per auth transition. A component redirects declaratively when its own state changes. ProtectedRoute owns "the user is gone". Logout doesn't navigate at all.
3. queryClient.clear() doesn't tell anyone
With the navigate() removed, logout called the API and then queryClient.clear(). The click did nothing. The page stayed exactly where it was.
clear() empties TanStack Query's cache map, but it doesn't notify the observers that are already mounted. The useQuery for /me inside the layout kept reporting its last value, a logged-in user, so ProtectedRoute had no reason to redirect. The fix is to invalidate instead:
async function logout() {
await authApi.logout(); // backend clears the cookies
await queryClient.invalidateQueries();
}Invalidation makes the active /me query refetch. It gets a real 401, the refresh also fails, the interceptor sets the user to null from outside the original call stack, and ProtectedRoute does the one redirect. Three pieces, each doing one job.
How verification works now
These bugs share a shape: correct types, plausible code, and broken behaviour that only shows up as a sequence of real requests in a real browser. So any change that touches the auth context, the API client, the route config or the permission gate has to be exercised live against the real backend before it's called done:
- Start the dev server and log in for real.
- Run the changed flow, then log out.
- Watch the console for repeating errors, which are the signature of a reload loop. A single
/meplus/refresh401 pair on a logged-out visit is expected and fine. - Check the network tab and compare response shapes against what the backend actually returns, not what the types assume.
The same habit caught a smaller gotcha: npx tsc --noEmit at the root checks nothing in a project-references setup. The real gate is tsc -b, which is what pnpm build runs.
Keeping the rewrite from rotting
A rewrite this size drifts back into a mess unless the structure is enforced. The code is split into two layers. src/core is the design system with zero app knowledge. src/shared is the app's infrastructure: the API client, auth, layout, and hooks that know the backend's shapes. The test for which side something belongs on is whether moving it to a different admin project would need even one line changed. ESLint's no-restricted-imports enforces that per folder: core/ can't import from shared/, configs/, modules/, i18n/ or store/, and one module domain can't reach into another's internals.
Since the migration the console has grown to nine domain groups and more than fifty feature folders, each with its own changelog. That growth is the part that makes me confident the call was right.
What I'd do again
- Measure the BFF before deleting it. "6.2% of the code" turned an argument about architecture into a plan with a cost attached.
- Same origin first, then cookies. With one origin, httpOnly cookies from the backend just work, and the frontend never touches a token.
- Send the CSRF token in the exact form the backend expects. Check which Spring Security handler it uses.
- Make refresh single-flight, with an allowlist of error codes, and never let refresh-failure handling reload the page.
- One redirect owner per auth transition. Logout clears state; the route guard redirects.
- Use
invalidateQueries()when mounted queries must react.clear()is for when nothing is watching. - Give nginx a resolver if the upstream can change IP.
- Treat green types as a starting point. Auth, data fetching and navigation get clicked through in a browser before they're done.
The full project, including the call-center and job-posting work that came after the migration, is in the Saramin Vietnam case study.
Related posts


