Fraud does not come in one shape. A carder tests forty stolen cards with tiny payments and then spends big on the three that pass. An account-takeover crew logs in from a new device, changes the password, changes the shipping address and orders a laptop. A promo abuser does nothing technically wrong at all, just with thirty accounts whose phone numbers differ by one digit. No single detection method catches all of that, and every vendor that says otherwise is selling a rule engine with a nicer dashboard.
So I built a platform that does not pick one. It runs five engines side by side (rules, supervised ML, unsupervised ML, a graph, and an LLM assistant for regulation) and turns them into one decision with reasons: approve, review or decline. It is open source under AGPL-3.0, and it is running.
The deck, if you prefer slides
Open the slides full screen ↗
Arrow keys move between slides, L switches between English and Bahasa Indonesia, and O shows the overview.
Five engines, because each one is blind somewhere
- Rules are the veteran analyst's checklist. Fast, transparent, and the only thing an auditor fully trusts. They catch what you already know about.
- Supervised ML learns from labelled cases and catches combinations of weak signals nobody would write as a rule. It is only as good as its labels.
- Unsupervised ML is the "something is off" detector: isolation forest, LOF and an autoencoder for anomalies, HDBSCAN and friends for clusters. It finds patterns nobody has labelled yet.
- The graph links customers through shared emails, phones, devices, IPs, cards, bank accounts and addresses, including similar ones. Fraud is a network, not a row. Two hops from a confirmed fraudster is a real signal.
- The LLM assistant reads regulations (OJK, BI, internal SOPs), checks whether existing rules still match them, and proposes new rules. It never activates anything by itself.
From raw event to decision
An event arrives through an API or a webhook in whatever shape the client already has. The platform maps it to a canonical event, persists it before any engine runs (ingest never loses an event), resolves graph entities, computes velocity and history features in SQL, then calls ML and the graph in parallel, and finally evaluates the rules with all of that as context. The target is under 150 ms at p95 without the LLM.
If the graph or ML service times out, it is dropped from the combination and the decision records that it ran degraded. If the rule service fails, the fallback is review, not approve. Failing open is how fraud teams get fired.
Why noisy-OR and not a weighted average
This is the design decision I like most, because the data made it for me. The first end-to-end simulation combined engine scores with a weighted average. In a project with no ML model yet, a strong rule hit of 60 was diluted to 45 by a graph engine that simply had nothing to say, and fell under the review threshold. Out of 111 fraud events, four were reviewed.
The default is now noisy-OR: score = 100 · (1 − Π(1 − sᵢ/100)^(wᵢ/w_max)). Independent evidence accumulates, and an engine at zero multiplies by one, so silence never drowns out a strong signal. A weighted average is still available as a project setting for teams that want a calibrated blend.
A rule language analysts actually own
The rule engine is a pure Rust crate with no IO. It supports five kinds of rule: simple (compare a field with a value or with another field), velocity (aggregate over a window with group-by, including z-score, gaussian tail, linear trend and poisson), composite (velocity over a filtered history), reference (white, black and watch lists created at runtime) and graph. Operands can be formulas such as F(x,y,z) = 2x + 2^y / z^2.
Every rule has three outcomes, not two: match, no match and trapped. Trapped means the rule could not be evaluated: a null field, a division by zero, too few samples for a statistic. It is logged with a reason and can be routed to review. A rule that silently returns "false" on missing data is a hole, not a rule.
Multi-tenant, and many projects per tenant
A marketplace does not have one fraud problem. Checkout, post-payment, returns, promo, account security and payouts behave differently and need different rules. In this platform each of them is a project with its own rules, models, graph, data sources and regulations, and a company (tenant) can have as many as it needs. Isolation between companies is enforced in the database with PostgreSQL Row-Level Security, which fails closed when no tenant is set.
Data comes in as it is. Upload a CSV, Excel or JSON sample, point it at a SQL table, or push through a webhook. The system infers the schema (it understands column names like tgl_transaksi, nominal and no_hp), suggests a mapping, hashes card and account numbers with a per-tenant pepper, and makes every column usable in rules and models with no code change.
Governance is a feature, not paperwork
- Maker–checker: rules, models, mappings and AI proposals need approval from a different person. Self-approval returns 403, and a database CHECK enforces it too.
- Shadow mode and backtest: a new rule can run on live traffic without affecting decisions, and you see its catches and false positives on history before you approve it.
- Immutable versions and an append-only audit log: every decision records
rule_id@version and the model versions, and a trigger blocks UPDATE and DELETE on the audit table.
- The LLM can only propose. Its single write tool creates a pending proposal. One human approval makes it a shadow rule; a second makes it active. The model runs locally on Ollama, so regulation documents and customer data never leave the server.
Does it work?
I trained on one simulated history and then scored a run with a new seed and new fraud rings, customers the models had never seen, against ground truth the platform never received. With default thresholds:
| Project | Recall | Precision | FPR | AUC |
| checkout | 95.2% | 40.0% | 7.4% | 0.986 |
| post-payment | 82.2% | 17.7% | 32.1% | 0.862 |
| returns | 100% | 13.7% | 34.7% | 0.998 |
| promo | 100% | 63.4% | 5.5% | 0.999 |
Ranking is strong everywhere (AUC 0.86 to 0.999). Recall is 100% for carding, system breach, promo abuse, refund abuse and bank account takeover, 84% for account takeover and 78% for money mules. The default threshold of 50 is clearly wrong for post-payment and returns, where it gives a third of legitimate events a review. That is the point of per-project thresholds: calibrated to a 5% false-positive rate, returns keeps 100% recall at 52% precision.
This is synthetic data. Absolute numbers will be different in production. What transfers is the method: train on history, score unseen traffic, join against truth the system never saw.
What I learned, honestly
- Labels are the bottleneck. Start with rules and anomaly detection, and let the case queue produce the labels that supervised ML needs.
- The graph needs confirmed fraud to be useful. Its AUC was about 0.5 on brand-new rings, which is expected: nobody connected to them had been labelled yet.
- Money mules are the hardest typology. Next up are fan-in/fan-out velocity rules on the recipient fingerprint and graph community features.
- AI proposes, humans decide. Interpreting regulation stays with compliance. A local LLM keeps data private but costs CPU, RAM and ideally a GPU.
Stack and running it yourself
Rust (axum, sqlx) for the latency-critical services: core API and orchestrator, rule service, graph service. Python 3.12 (FastAPI, PyTorch, scikit-learn, pandas) for ML, the LLM service and ingest. PostgreSQL 16 with one schema per service, OpenSearch for per-tenant vectors, Ollama for local models, Nuxt 4 for the UI, and Traefik as the only exposed service.
The whole stack starts with docker compose up -d --build on a machine with about 16 GB of RAM. The repository has a quickstart, a non-technical overview for management and compliance, user and developer guides, and the evaluation method so you can reproduce the numbers above.
Code, issues and pull requests: github.com/situkangsayur/fraud_detection_engine.
Fraud tidak datang dalam satu bentuk. Pelaku carding menguji empat puluh kartu curian dengan transaksi kecil, lalu belanja besar dengan tiga kartu yang lolos. Pelaku account takeover login dari perangkat baru, mengganti password, mengganti alamat kirim, lalu memesan laptop. Pelaku abuse promo bahkan tidak melanggar apa pun secara teknis, hanya saja ia punya tiga puluh akun dengan nomor HP yang berbeda satu digit. Tidak ada satu metode deteksi yang bisa menangkap semua itu, dan vendor yang bilang sebaliknya biasanya sedang menjual rule engine dengan dashboard yang lebih cantik.
Jadi saya membangun platform yang tidak memilih salah satu. Platform ini menjalankan lima mesin berdampingan (rule, ML terawasi, ML tak terawasi, graph, dan asisten LLM untuk regulasi) lalu menggabungkannya menjadi satu keputusan beserta alasannya: approve, review, atau decline. Kodenya open source dengan lisensi AGPL-3.0, dan sudah berjalan.
Versi slide, kalau lebih suka
Buka slide layar penuh ↗
Tombol panah untuk pindah slide, L untuk ganti bahasa Inggris/Indonesia, dan O untuk melihat semua slide.
Lima mesin, karena masing-masing punya titik buta
- Rule adalah checklist analis senior. Cepat, transparan, dan satu-satunya yang benar-benar dipercaya auditor. Menangkap pola yang sudah kita kenal.
- ML terawasi belajar dari kasus berlabel dan menangkap kombinasi sinyal lemah yang tidak akan pernah ditulis orang sebagai rule. Kualitasnya sebaik labelnya.
- ML tak terawasi adalah detektor "ada yang aneh": isolation forest, LOF, dan autoencoder untuk anomali, HDBSCAN dan kawan-kawan untuk klaster. Menemukan pola yang belum pernah diberi label.
- Graph menghubungkan pelanggan lewat email, HP, perangkat, IP, kartu, rekening, dan alamat yang sama, termasuk yang mirip. Fraud itu jaringan, bukan satu baris data. Berjarak dua langkah dari penipu yang terkonfirmasi adalah sinyal nyata.
- Asisten LLM membaca regulasi (OJK, BI, SOP internal), mengecek apakah rule yang ada masih sesuai, dan mengusulkan rule baru. Ia tidak pernah mengaktifkan apa pun sendiri.
Dari event mentah ke keputusan
Event masuk lewat API atau webhook dalam format apa pun yang sudah dimiliki klien. Platform memetakannya ke event kanonik, menyimpannya sebelum mesin mana pun berjalan (ingest tidak pernah kehilangan event), meresolusi entitas graph, menghitung fitur velocity dan histori di SQL, memanggil ML dan graph secara paralel, lalu mengevaluasi rule dengan semua itu sebagai konteks. Targetnya di bawah 150 ms pada p95 tanpa LLM.
Kalau service graph atau ML timeout, ia dikeluarkan dari kombinasi dan keputusan mencatat bahwa ia berjalan dalam mode degradasi. Kalau rule service gagal, fallback-nya review, bukan approve. Gagal-terbuka adalah cara tercepat tim fraud kehilangan pekerjaan.
Kenapa noisy-OR, bukan rata-rata berbobot
Ini keputusan desain yang paling saya suka, karena datanya yang memutuskan. Simulasi end-to-end pertama menggabungkan skor mesin dengan rata-rata berbobot. Di project yang belum punya model ML, rule hit yang kuat dengan skor 60 terencerkan menjadi 45 oleh mesin graph yang memang belum punya apa-apa untuk dikatakan, dan jatuh di bawah threshold review. Dari 111 event fraud, hanya empat yang masuk review.
Default-nya sekarang noisy-OR: skor = 100 · (1 − Π(1 − sᵢ/100)^(wᵢ/w_max)). Bukti independen terakumulasi, dan mesin bernilai nol dikali satu, jadi mesin yang diam tidak pernah menenggelamkan sinyal yang kuat. Rata-rata berbobot tetap tersedia sebagai setting project untuk tim yang ingin campuran yang terkalibrasi.
Bahasa rule yang benar-benar dikendalikan analis
Rule engine-nya adalah crate Rust murni tanpa IO. Ada lima jenis rule: simple (bandingkan field dengan nilai atau dengan field lain), velocity (agregasi dalam window dengan group by, termasuk z-score, gaussian tail, tren linear, dan poisson), composite (velocity atas histori yang difilter), reference (daftar putih, hitam, dan pantau yang dibuat saat runtime), dan graph. Operand bisa berupa formula seperti F(x,y,z) = 2x + 2^y / z^2.
Setiap rule punya tiga hasil, bukan dua: match, no match, dan trapped. Trapped berarti rule tidak bisa dievaluasi: field kosong, pembagian dengan nol, sampel terlalu sedikit untuk statistik. Hasil ini dicatat beserta alasannya dan bisa diarahkan ke review. Rule yang diam-diam mengembalikan "false" saat datanya kosong itu lubang, bukan rule.
Multi-tenant, dan banyak project per tenant
Sebuah marketplace tidak punya satu masalah fraud. Checkout, pasca-pembayaran, retur, promo, keamanan akun, dan payout punya perilaku berbeda dan butuh rule berbeda. Di platform ini masing-masing adalah project dengan rule, model, graph, sumber data, dan regulasinya sendiri, dan satu perusahaan (tenant) bisa punya sebanyak yang dibutuhkan. Isolasi antar perusahaan dijamin di database dengan PostgreSQL Row-Level Security, yang menolak akses bila tenant tidak di-set.
Data masuk apa adanya. Unggah contoh CSV, Excel, atau JSON, arahkan ke tabel SQL, atau kirim lewat webhook. Sistem mengenali skemanya (termasuk nama kolom seperti tgl_transaksi, nominal, dan no_hp), menyarankan mapping, meng-hash nomor kartu dan rekening dengan pepper per tenant, dan membuat semua kolom langsung bisa dipakai di rule dan model tanpa mengubah kode.
Tata kelola adalah fitur, bukan sekadar dokumen
- Empat mata (maker–checker): rule, model, mapping, dan usulan AI harus disetujui orang lain. Menyetujui sendiri menghasilkan 403, dan CHECK di database juga menjaganya.
- Mode bayangan dan backtest: rule baru bisa dijalankan di trafik nyata tanpa memengaruhi keputusan, dan efeknya pada data historis (yang tertangkap dan yang salah tangkap) terlihat sebelum disetujui.
- Versi immutable dan audit log append-only: setiap keputusan mencatat
rule_id@version dan versi model, dan trigger memblokir UPDATE dan DELETE pada tabel audit.
- LLM hanya bisa mengusulkan. Satu-satunya tool tulisnya membuat usulan pending. Satu persetujuan manusia menjadikannya rule shadow; persetujuan kedua menjadikannya aktif. Modelnya berjalan lokal di Ollama, jadi dokumen regulasi dan data pelanggan tidak pernah keluar dari server.
Apakah berhasil?
Saya melatih model pada satu histori simulasi, lalu menilai run dengan seed baru dan ring fraud baru, yaitu pelanggan yang belum pernah dilihat model, dan membandingkannya dengan ground truth yang tidak pernah dikirim ke platform. Dengan threshold default:
| Project | Recall | Precision | FPR | AUC |
| checkout | 95,2% | 40,0% | 7,4% | 0,986 |
| post-payment | 82,2% | 17,7% | 32,1% | 0,862 |
| returns | 100% | 13,7% | 34,7% | 0,998 |
| promo | 100% | 63,4% | 5,5% | 0,999 |
Kualitas ranking tinggi di semua project (AUC 0,86 sampai 0,999). Recall 100% untuk carding, sistem dibobol, abuse promo, abuse retur, dan pengambilalihan rekening; 84% untuk account takeover dan 78% untuk money mule. Threshold default 50 jelas tidak cocok untuk post-payment dan retur, karena sepertiga event yang sah ikut masuk review. Itulah gunanya threshold per project: dikalibrasi ke false-positive rate 5%, project retur tetap punya recall 100% dengan precision 52%.
Ini data sintetis. Angka absolutnya akan berbeda di produksi. Yang bisa dibawa adalah metodenya: latih pada histori, nilai trafik yang belum pernah dilihat, bandingkan dengan kebenaran yang tidak pernah dilihat sistem.
Yang saya pelajari, apa adanya
- Label adalah bottleneck. Mulai dengan rule dan deteksi anomali, dan biarkan antrean case menghasilkan label yang dibutuhkan ML terawasi.
- Graph butuh fraud yang terkonfirmasi. AUC-nya sekitar 0,5 pada ring yang benar-benar baru, dan itu wajar: belum ada pelanggan terkait yang diberi label.
- Money mule adalah tipologi paling sulit. Berikutnya: rule velocity fan-in/fan-out pada sidik jari penerima dan fitur komunitas graph.
- AI mengusulkan, manusia memutuskan. Interpretasi regulasi tetap di tim compliance. LLM lokal menjaga privasi data, tapi butuh CPU, RAM, dan idealnya GPU.
Stack dan cara menjalankannya
Rust (axum, sqlx) untuk service yang kritis terhadap latensi: core API dan orkestrator, rule service, graph service. Python 3.12 (FastAPI, PyTorch, scikit-learn, pandas) untuk ML, service LLM, dan ingest. PostgreSQL 16 dengan satu skema per service, OpenSearch untuk vektor per tenant, Ollama untuk model lokal, Nuxt 4 untuk UI, dan Traefik sebagai satu-satunya service yang terekspos.
Seluruh stack jalan dengan docker compose up -d --build di mesin dengan RAM sekitar 16 GB. Repositorinya berisi quickstart, ringkasan non-teknis untuk manajemen dan compliance, panduan pengguna dan developer, serta metode evaluasi supaya angka di atas bisa direproduksi.
Kode, issue, dan pull request: github.com/situkangsayur/fraud_detection_engine.