Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Day 1: Build the Foundation — Ingest, Search, Retrieve
Module 1: The Legal Engineer’s Landscape
- Learning objectives — understand the role of a legal engineer, where AI fits into legal work, and the two critical risks that permeate the field.
- Topics
-
- The legal-engineer role and current market demand.
- AI applications: eDiscovery, review, contracts, research, investigations; the EDRM model explained simply.
- Build vs. buy considerations.
- The two pervasive risks: confidentiality/privilege and defensibility.
Module 2: Legal Data Is Messy — Ingestion and Extraction
- Learning objectives — handle the reality of processing legal data at scale.
- Topics
- Handling 1,400+ file types, emails, PSTs, scanned papers, load files (.dat/.opt); understanding relevant embedded metadata.
- Text extraction (Tika), OCR, and de-duplication strategies.
- Lab: FreeEed Ingestion — build an ingestion pipeline over a deliberately messy document set (email/PST, scans, load files).
Module 3: Search and Retrieval — the Foundation
- Learning objectives — build the core eDiscovery primitive: the ability to find anything inside everything.
- Topics — full-text search and indexing (Solr/Lucene); relevance, metadata, and date filtering; searching across OCR’d content.
- Lab: eDiscovery Search — index a corpus and run real eDiscovery-style searches, including within OCR’d scans.
Module 4: RAG for Legal Documents — with Citations
- Learning objectives — build a Retrieval-Augmented Generation (RAG) system over legal documents that cites its sources.
- Topics
- Why retrieval, not fine-tuning, is preferred for sensitive material—the model never ingests the documents directly.
- Chunking, embeddings, and critically, citations / provenance.
- Multi-document and thread summarization.
- Lab: Legal RAG with Citations — build a RAG Q&A system over a document set that answers queries with source citations.
Day 2: Make It Private, Defensible, and Shippable
Module 5: Privacy, Privilege, and Local Serving — the Privilege Trap
- Learning objectives — keep legal data local and able to certify its security.
- Topics
- What happens to data when it hits a cloud AI service.
- Privilege waiver, duty of competence, and the spectrum of "private" (contractual vs. physical).
- The precedent of Morgan v. V2X and why local solutions are court-defensible.
- Serving local models (Ollama / vLLM) and monitoring outbound traffic.
- Lab: Local Model + Egress Proof — run a local model end-to-end and prove via monitoring that no data exited the environment.
Module 6: Defensible AI Review
- Learning objectives — measure and document an AI review process so it holds up in court.
- Topics
- Critical metrics for court: recall, elusion, precision, ground-truth validation; TAR / active learning.
- Transparency (why did it code this document?) and reproducibility — pin the model version, fix settings, log everything.
- The "defensible case snapshot" that allows someone to re-run your review a year later and achieve identical results.
- Lab: Defensible Review — measure an AI review against a blind ground truth and produce a reproducibility bundle.
Module 7: Ship It — Workflow, Private Deployment, and Governance
- Learning objectives — assemble components into a workflow, deploy privately, and score the system.
- Topics
- A multi-step legal workflow (ingest → search → summarize → review → produce) with human-in-the-loop oversight.
- Essentials for private/on-premises deployment (containerization; keeping data on-site).
- Brief AI governance for legal contexts and scoring the system using SAIS-100 (the Elephant Scale Secure AI Score).
- Lab: Score and Package — wire a multi-step workflow, score it with SAIS-100, and package it for private deployment.
Capstone (integrated across Day 2)
- Construct a private, defensible legal-AI application end-to-end — ingest a messy corpus, search it, answer questions with citations using a local model, measure a defensible review, and package for private deployment.
- Participants leave with a portfolio project that mirrors the responsibilities of a legal engineer role.
Optional Day 3 / Advanced Modules (deliverable as a 3rd day or a modular series)
- Investigations: Entities, Relationships, and Timelines — extract people/orgs/dates, reconstruct email threads, build chronologies, map near-duplicates and document lineage. Lab: build a timeline and entity/relationship view.
- Agentic and Multi-Step Legal Workflows (deep dive) — richer orchestration, contract analysis, multi-document synthesis, tool use, and guardrails as a design principle. Lab: build a multi-step workflow with a human checkpoint.
- Deployment at Scale — on-premises and appliance deployment, distributed processing for large volumes, regulated environments (CJIS, government, higher-ed), hardware sizing. Lab: containerize and scale a processing job across workers.
- Governance and Compliance Deep-Dive — the AI-regulation landscape (100+ US state AI laws, the EU AI Act), audit requirements, and a full SAIS-100 governance audit. Lab: audit a legal-AI system against a governance/defensibility checklist.
Requirements
- Familiarity with Python and basic APIs.
- Helpful: User-level understanding of Large Language Models (LLMs) — no machine learning background is required as the course builds necessary mental models.
- No prior legal background is required; essential legal concepts are taught within relevant contexts.
Audience
- Software and AI engineers transitioning into legal technology.
- Engineers at legal-tech companies who need deeper domain expertise in legal processes.
- Technically inclined legal, eDiscovery, or information governance professionals interested in building tools rather than merely purchasing them.
- Individuals aiming for "legal engineer" or "AI legal engineer" roles.
14 Hours
Testimonials (1)
That i gained a knowledge regarding streamlit library from python and for sure i'll try to use it to improve applications in my team which are made in R shiny