Skip to content
Early Beta — internal transactions recorded, seeding independent demand. See the numbers
All listings
Machine Learning Data DatasetPlatform-seededFree

YES, IT'S VERMARCABLE

Machine Learning: Core Concepts & Model-Architecture Glossary

A free, platform-curated glossary of the core machine-learning concepts, architectures, and training methods practitioners rely on, with precise plain-language definitions consistent with the published research literature and NIST terminology. Each record gives the term, its definition, a category, and a short reference note. Ideal for onboarding, technical writing, agent grounding, and buyers who need a shared, citable model vocabulary.

0 sold 466 views7/14/2026
Free preview
{
  "_type": "curated_open_data",
  "as_of": "2026-07",
  "links": {
    "canonical": "https://verticalmarketplace.ai",
    "docs_for_llms": "https://verticalmarketplace.ai/llms.txt",
    "sell_your_own": "https://verticalmarketplace.ai/api/marketplace/listings",
    "vertical_listings": "https://verticalmarketplace.ai/api/marketplace/listings?vertical=ai-ml"
  },
  "records": [
    {
      "term": "Supervised learning",
      "category": "training",
      "reference": "Foundational ML paradigm",
      "definition": "Training a model on input data paired with known labels so it can predict labels for new inputs"
    },
    {
      "term": "Unsupervised learning",
      "category": "training",
      "reference": "Includes clustering and dimensionality reduction",
      "definition": "Finding structure or patterns in data that has no labels"
    },
    {
      "term": "Neural network",
      "category": "architecture",
      "reference": "Basis of deep learning",
      "definition": "A layered model of interconnected weighted units that approximates complex functions"
    },
    {
      "term": "Transformer",
      "category": "architecture",
      "reference": "Introduced in 'Attention Is All You Need' (2017)",
      "definition": "An architecture that uses self-attention to weigh relationships across an entire input sequence"
    },
    {
      "term": "Parameter",
      "category": "architecture",
      "reference": "Model size is often reported as parameter count",
      "definition": "A learned numerical weight adjusted during training that stores what a model has learned"
    },
    {
      "term": "Token",
      "category": "architecture",
      "reference": "Inputs and outputs are measured in tokens",
      "definition": "A unit of text (word piece or character) that a language model processes"
    },
    {
      "term": "Fine-tuning",
      "category": "training",
      "reference": "Adaptation method",
      "definition": "Further training a pretrained model on task-specific data to specialize it"
    },
    {
      "term": "Inference",
      "category": "deployment",
      "reference": "The production phase after training",
      "definition": "Running a trained model on new inputs to produce predictions or outputs"
    },
    {
      "term": "Overfitting",
      "category": "evaluation",
      "reference": "Countered with regularization and validation",
      "definition": "When a model memorizes training data and generalizes poorly to new data"
    },
    {
      "term": "Gradient descent",
      "category": "training",
      "reference": "Core training algorithm",
      "definition": "An optimization method that iteratively adjusts weights to reduce a loss function"
    },
    {
      "term": "Embedding",
      "category": "representation",
      "reference": "Enables semantic search and retrieval",
      "definition": "A dense numeric vector that represents an item so similar items sit close together"
    },
    {
      "term": "Large language model (LLM)",
      "category": "architecture",
      "reference": "Class of foundation models",
      "definition": "A transformer trained on very large text corpora to generate and understand language"
    },
    {
      "term": "Reinforcement learning from human feedback (RLHF)",
      "category": "training",
      "reference": "Used to align conversational assistants",
      "definition": "Aligning model behavior using a reward model trained on human preference data"
    },
    {
      "term": "Retrieval-augmented generation (RAG)",
      "category": "deployment",
      "reference": "Reduces reliance on parametric memory",
      "definition": "Supplying a model with retrieved documents at inference time to ground its answers"
    },
    {
      "term": "Hallucination",
      "category": "evaluation",
      "reference": "Key reliability risk",
      "definition": "When a model generates fluent content that is factually incorrect or unsupported"
    }
  ],
  "sources": [
    {
      "url": "https://www.nist.gov/itl/ai-risk-management-framework",
      "name": "NIST AI Risk Management Framework (terminology)"
    },
    {
      "url": "https://arxiv.org/",
      "name": "arXiv preprint server (Cornell University)"
    }
  ],
  "category": "glossary",
  "vertical": "ai-ml",
  "data_note": "All records are public-domain facts compiled from the cited sources as of the asOf date. This content is authored and served by the platform itself — it is not seller data, so the marketplace's zero-storage promise about seller datasets is unaffected.",
  "record_count": 15,
  "what_this_is": "A platform-published open-data listing curated by Open Data Desk, the marketplace's in-house public-data seller. It is real free inventory: it counts in marketplace statistics and is purchasable for $0 through the normal purchase flow, which delivers this payload with an Ed25519-signed receipt.",
  "record_schema": {
    "term": "The concept, method, or architecture",
    "category": "Grouping such as training, architecture, or evaluation",
    "reference": "Origin, standard, or clarifying note",
    "definition": "Plain-language meaning"
  },
  "buyer_use_cases": [
    "Standardize model terminology across product, legal, and engineering teams",
    "Ground a chatbot or agent with precise, citable ML definitions",
    "Speed up onboarding and technical documentation with a ready glossary"
  ]
}

The full dataset is delivered after purchase. Fingerprint: sha256:b59b70c7b3b46635d55b46fdaf0677b8d8fcaa9b41790f74b963305f2e6b44e0

Questions & answers

No answered questions yet — ask the seller anything about this listing.

License — Vertical Marketplace Data License v1

Data is contributed by independent third-party sellers. Vertical Marketplace facilitates the transaction; sellers keep 95% on everyday sales from $20 to $49,999.99 under the year-one founding rate locked through 2027-06-30 (full schedule: GET /api/meta). Prohibited content (digital keys/licenses/game codes, and health data the seller does not own — e.g. patient records) is not permitted; individuals may sell their own personal health data only via the signed Health Data Consent Flow. See /terms.

Permitted
  • Use the purchased data for your own commercial and non-commercial work
  • Create derivative analyses, models, and works from the data
Restricted
  • No reselling or re-listing the purchased data on this or any other marketplace
  • No redistributing the raw data payload as-is to third parties
  • Exclusive listings are sold to a single buyer and delisted on purchase
  • Limited listings are sold to a capped number of buyers and delisted once sold out
Price
FREE
Use an existing account (optional)

Use the account that owns your wallet, saved card or purchase. You can also continue as a guest without a key.

For agents — buy by prompt

Bring your own agent: open HTTP API + MCP, any compatible runtime. Compatibility and vendor independence.

Seller
S-000002
Platform-affiliated seller — common ownership
No ratings yet
First Mover