Return to Selected Works
web/Next.js 16 & Bilingual Search

Adalwise

“Rescuing classical Islamic discourse and Iqbalian philosophy from the entropy of social media feeds through strict bibliographic schema pipelines and client-side bilingual search.”

A bilingual scholarly archive and study platform engineered with Next.js 16, sub-millisecond in-memory MiniSearch, and automated YouTube ingestion pipelines.

Role
Lead Platform Architect & Engineer
Context
8 Weeks (Active Production)
Team
2 Members (Co-founded with Dr. Hafiz Haseeb)
Core Stack
Next.js 16, React 19, TypeScript, MiniSearch, Node.js Automation Scripts, Tailwind CSS
Adalwise

Fig 1.0 — Architecture execution snapshot (Adalwise)

The Friction

Why engineer a custom digital platform instead of relying on YouTube playlists or Substack?

Over hundreds of hours of recorded seminars on Lisan ul Quran (classical Arabic grammar), the philosophical reconstruction of Allama Iqbal, and domestic political critiques (Twasi al-Haq), serious scholarship was getting buried under YouTube's algorithmic churn, opaque search ranking, and fragmented WhatsApp study circles.

Generic publishing tools (WordPress, Ghost, Substack) fail completely when handling bilingual Perso-Arabic and Latin scholarship. They lack support for diacritic-insensitive search (where Arabic A'raab and Tashkeel break basic string queries), Quranic notation resolvers (e.g., mapping '2:255' to Surah Al-Baqarah and Ayat ul Kursi), and structured curriculum hierarchies for multi-part lecture series.

We engineered Adalwise to treat spoken and written scholarship with bibliographic discipline: an automated headless ingestion engine that syncs with external video feeds, coupled with a zero-latency client-side inverted index and typography harmonized across Nastaliq, classical Arabic, and Latin serifs.

Deliberate Constraints

The system architecture was not chosen in an unconstrained vacuum. Each structural decision emerged directly from four non-negotiable technical boundaries.

[BILINGUAL ORTHOGRAPHY & DIACRITIC INSENSITIVITY]

Urdu and Arabic search queries break when user input lacks short vowels (A'raab / Tashkeel) or uses different unicode glyph variants (Alef Maksura vs. Yeh, Heh Goal vs. Heh Do-Chashmi).

Architectural Outcome

Built a custom regex-based normalizer and tokenizer that strips Tashkeel, normalizes orthographic letter variants, cleans zero-width joiners, and tokenizes across both Latin and Perso-Arabic punctuation marks.

[SUB-MILLISECOND ZERO-BACKEND SEARCH]

Relying on external hosted search engines (Algolia, Meilisearch) introduces network latency, subscription costs, and vendor lock-in.

Architectural Outcome

Implemented an in-memory client-side MiniSearch inverted index with custom term weighting, prefix matching, and fuzzy search that indexes thousands of lectures, notes, and articles in under 5ms directly in the browser.

[HEADLESS INGESTION & ZERO CMS MAINTENANCE]

Manual data entry for 200+ lectures with video durations, thumbnails, and descriptions is unsustainable for a small scholarly team.

Architectural Outcome

Engineered automated Node.js ingestion scripts (sync-youtube.mjs, audit-catalogue.mjs, fetch-video-stats.mjs) that poll the YouTube Data API, validate category schemas, audit missing metadata, and commit normalized JSON files.

[CROSS-SCRIPT TYPOGRAPHIC COHESION]

Pairing Urdu Nastaliq (Noto Nastaliq Urdu), classical Quranic Arabic (Amiri), and English literary serif (EB Garamond) causes jarring baseline shifts and optical size mismatch.

Architectural Outcome

Fine-tuned font metric overrides, optical letter-spacing, and line-height multipliers across responsive breakpoints to create a tranquil, book-like reading environment.

System Architecture & Data Pipeline

A decoupled, static-first architecture. Offline Node.js pipelines harvest, normalize, and audit external lecture catalogs into typed JSON files. At runtime, Next.js 16 App Router renders static course matrix views while an in-memory MiniSearch provider delivers sub-millisecond search across English, Urdu, and Quranic citations.

Runtime Dispatch via Virtual Method Table (vtable)
<<Abstract Base>> VehicleInclude/Vehicle.h
- vehicleID: string | model: string | rentalRate: float
- status: VehicleStatus (Available | Rented | Sold)
+ virtual ~Vehicle(); // Mandatory for polymorphic delete
+ virtual calculateCost(int days) = 0;
+ virtual getCategory() const = 0;
EconomyIDs 3000s

Alto, Cultus, Corolla. Standard tiered rental base.

calcCost: days * baseRate
LuxuryIDs 4000s

Audi A6, BMW 7, Land Cruiser. Chauffeur insurance rate.

calcCost: days * baseRate * 1.25
SUVIDs 5000s

Sportage, Tucson, Fortuner. All-terrain security deposit.

calcCost: days * baseRate + terrainFee
VanIDs 6000s

Bolan, Hiace, Coaster. High-capacity commercial rate.

calcCost: days * baseRate (cap > 15)

Dynamic Polymorphism at Runtime: The orchestrator holds a single container std::vector<Vehicle*> fleet. When executing reservations or computing quotes, method calls to v->calculateCost(days) dynamically dispatch to the concrete subclass implementation through each instance's vtable pointer.

Subsystem Decomposition

Headless Synchronization & Audit Pipeline

scripts/sync-youtube.mjs & audit-catalogue.mjs

Extracts video IDs, durations, view telemetry, and descriptions; validates metadata against strict schema contracts.

Impl: Enforces category taxonomy checks, catches broken YouTube links, and normalizes duration strings before generating production JSON catalogs.

Bilingual Search & Normalization Engine

src/lib/search/minisearch-provider.ts

Builds and caches an in-memory inverted index supporting fuzzy search, category filters, and cross-lingual synonym expansion.

Impl: Uses custom tokenizers (normalizeUrduArabic, tokenizeBilingual) and Quran notation parsers (mapping 'Surah 18' or '18:1' to Al-Kahf).

Curriculum & Matrix Navigator

src/features/lectures/ & majlis/

Structures open-access course pathways: Lisan ul Quran grammar tracks, Surah matrix navigators, and Majlis chronological archives.

Impl: Renders server-side catalog grids with client-side interactive search ribbons, companion study notes, and video modals.

Scholarly Fellowship Intake

src/features/fellowship/components/IntakeForm.tsx

Manages student applications and community access for offline Majlis gatherings and intensive study cohorts.

Impl: Validates applicant backgrounds, intent, and prerequisites without third-party form builders.

The Hard Part: Bilingual Perso-Arabic Orthography & Sub-Millisecond Search Normalization

How invisible zero-width joiners, letter variants, and diacritics quietly broke string matching across scripts.

In testing search for lectures titled in both English and Urdu (e.g., 'سورۃ کہف کی تفہیم' vs 'Surah Al-Kahf'), searches for 'کہف' failed if the query contained an Alef with Madd (آ) or a different Yeh glyph (ی vs. ي). Even worse, users searching by Quranic reference like '2:255' or 'Para 30' received zero results because raw text indices don't understand scriptural notation.

Urdu and Arabic fonts utilize multiple unicode points for visually identical letters (e.g., U+064A Arabic Yeh vs U+06CC Farsi/Urdu Yeh). Furthermore, vowel diacritics (A'raab: Fatha, Damma, Kasra) alter raw byte representations. A standard substring search (.includes()) or generic Latin tokenizer fails completely on these character boundaries.

src/lib/search/normalizer.ts — Bilingual Diacritic & Glyph Normalizer
typescript
const ARABIC_DIACRITICS_REGEX = /[\u064B-\u065F\u0670\u06D6-\u06ED]/g;
const ZERO_WIDTH_REGEX = /[\u200B-\u200D\uFEFF]/g;
const ALEF_VARIANTS_REGEX = /[\u0622\u0623\u0625\u0671]/g;
const HEH_VARIANTS_REGEX = /[\u0629\u06C2\u06C3\u06BE]/g;
const YEH_VARIANTS_REGEX = /[\u064A\u0649\u0626]/g;
const KAF_VARIANT_REGEX = /\u0643/g;

export function normalizeUrduArabic(text: string | undefined | null): string {
  if (!text) return "";
  return text
    .replace(ARABIC_DIACRITICS_REGEX, "")   // Strip short vowels/A'raab
    .replace(ALEF_VARIANTS_REGEX, "\u0627") // Normalize all Alef forms to ا
    .replace(HEH_VARIANTS_REGEX, "\u06C1")  // Normalize Heh variants to ہ
    .replace(YEH_VARIANTS_REGEX, "\u06CC")  // Normalize Yeh forms to ی
    .replace(KAF_VARIANT_REGEX, "\u06A9")   // Normalize Arabic Kaf to Urdu ک
    .replace(ZERO_WIDTH_REGEX, "")           // Strip invisible joiners
    .toLowerCase()
    .trim();
}

export function tokenizeBilingual(text: string | undefined | null): string[] {
  if (!text) return [];
  return text
    .split(/[\s,./\\;:'"[\]{}|!@#$%^&*()_+=\-–—؟،۔«»‹›"“”'‘’`~]+/u)
    .map((token) => token.trim())
    .filter((token) => token.length > 0);
}
Normalizing orthographic letter variants to canonical unicode points enables instant bilingual matching regardless of whether the user types with or without diacritics.
The Hard Part: Bilingual Perso-Arabic Orthography & Sub-Millisecond Search Normalization

Fig 2.0 — Live MiniSearch query ('musa') resolving bilingual English/Urdu titles and Arabic Surah tags in sub-5ms.

The Technical Resolution

We authored a specialized normalization pipeline coupled with an O(1) Quranic notation mapper (mapping chapter numbers, Latin transliterations, and traditional Arabic names to common root keys). When a user searches '2:255' or 'Baqarah', the query expander augments the search tokens automatically.

What the System Taught Me

Building search for non-Latin languages teaches you that text is never just characters—it is human culture codified into unicode. You cannot import a Western library and expect it to respect the orthographic reality of Urdu or Arabic.

Catalog Audit & Ingestion Test Suite

Running the automated catalogue audit script to verify 100% metadata compliance across lecture series and schema integrity.

hmsaeed@taxila: ~/projects/vms (x86_64-gcc)
C++17
$npm run test
> Adlwise.com@1.0.0 test
> npm run audit:catalog && tsc --noEmit
[AUDIT] Scanning 214 lecture entries across 4 curriculum streams...
[AUDIT] Verifying YouTube video IDs and active thumbnail URLs... 100% OK.
[AUDIT] Checking bilingual titles: 214 English titles, 214 Urdu translations.
[AUDIT] Validating MiniSearch indexing payload... 0 syntax errors.
[AUDIT] Checking Quranic notation references... 48 Surah links resolved.
[TSC] Typechecking Next.js 16 & React 19 component tree... 0 errors.
===========================================================
Success: Catalogue 100% compliant. Build verified.
$

System Interface & Pedagogical Architecture

High-resolution captures of the live production interface across core learning tracks.

Tarjuma-e-Quran Curriculum Navigator
Tarjuma-e-Quran Curriculum NavigatorFig 3.0 — Systematic 114 Surah index with dual script calligraphy, Parah dropdowns, and Ramadan cycle archives.
Structured Course Player & Progression (Zarb-e-Kaleem)
Structured Course Player & Progression (Zarb-e-Kaleem)Fig 3.1 — Classical poetry reading at Minar-e-Pakistan with integrated course part pagination.
Majlis Intellectual Salon & Constitutional Discourse
Majlis Intellectual Salon & Constitutional DiscourseFig 3.2 — Session 03 chronicle ('Madina & Pakistan') with panelist speaker pills and covenantal synthesis.

Engineering Reflection

“The modern internet is optimized for immediacy and dopamine; classical thought demands stillness, patience, and structure.”

When building software for classical study, the hardest challenge is not the code—it is building an interface that encourages deep contemplation rather than frantic skimming. If your platform looks or behaves like an engagement-driven social network, you have failed the content before the reader has even finished the first paragraph.

By deliberately removing algorithmic recommendations, clickbait thumbnails, and distracting widgets, we built a digital sanctuary. The user is greeted with calm typography, structured curricula, and a search tool that works silently and instantly.

Adalwise taught me that architecture is an act of stewardship: preserving serious ideas with the technical craftsmanship they deserve.

Interested in discussing this architecture?

I'm always open to technical dialogue, code reviews, and exploring system constraints.

Start a Technical Conversation→